Benchmarking LLMs for Real-World Human Behavior Simulation

Explore OmniBehavior, a benchmark using real-world data to evaluate LLMs on long-term, cross-scenario human behavior simulation and address model biases.

Effective Divergence Measures for Training GFlowNets

Explore optimized divergence measures to enhance GFlowNets training, improving convergence speed and sampling accuracy in generative modeling.

Strategic Algorithmic Monoculture in AI Coordination Games

Explore experimental insights on strategic algorithmic monoculture and AI coordination, revealing how agents adapt actions in multi-agent environments.

Process Reward Agents for Enhanced Knowledge-Intensive AI Reasoning

Discover how Process Reward Agents improve AI reasoning accuracy in knowledge-intensive tasks without retraining, boosting performance in medical benchmark...

E3-TIR: Boosting Tool-Integrated Reasoning Efficiency

Discover E3-TIR, a novel training paradigm enhancing tool-integrated reasoning with expert guidance and self-exploration for better AI performance.

Popular

Subscribe