How Much LLM Power Does a Self-Revising Agent Need?

Explore how much large language model (LLM) capacity self-revising agents require for optimal planning, reflection, and decision-making in AI systems.

T-STAR: Tree-Based Policy Optimization for Multi-Turn Agents

Discover T-STAR, a novel tree-structured framework enhancing multi-turn agent policy optimization through self-rectification and grafting techniques.

EVGeoQA: Benchmarking LLMs for Dynamic Geo-Spatial Tasks

Explore EVGeoQA, a benchmark evaluating LLMs on dynamic, multi-objective geo-spatial exploration, focusing on EV charging and real-time location data.

Planning Task Shielding: Detect and Repair Flaws in AI Planning

Discover how planning task shielding detects and repairs flaws in AI planning tasks by making them unsolvable to prevent errors and enhance reliability.

A-MBER: Benchmark for Long-Term Emotion Recognition AI

Discover A-MBER, a benchmark evaluating AI's ability to recognize emotions using long-term memory for more empathetic user interactions.

Popular

Subscribe