SEA-Eval: Benchmark for Evaluating Self-Evolving AI Agents

Discover SEA-Eval, a benchmark designed to evaluate self-evolving AI agents beyond episodic tasks, focusing on long-term performance and reliability.

PilotBench: Benchmarking AI Safety in General Aviation

Discover PilotBench, a benchmark evaluating AI models on safety and precision in general aviation flight predictions with real-world data.

Boost LLM Problem Solving with Tutor-Student Agents

Discover how tutor-student multi-agent interaction enhances Large Language Model problem solving with efficient, role-based collaboration.

StaRPO: Enhancing RL with Stability for Better Reasoning

StaRPO improves reinforcement learning by adding stability metrics, boosting logical consistency and accuracy in AI reasoning tasks.

SPPO: Efficient PPO for Long-Horizon Reasoning Tasks

Discover SPPO, a scalable Sequence-Level PPO algorithm improving stability and efficiency in long-horizon reasoning for Large Language Models.

Popular

Subscribe