Agentic-MME: Benchmarking Multimodal Agentic Intelligence

Discover Agentic-MME, a benchmark evaluating multimodal agentic capabilities with real-world tasks, stepwise checkpoints, and unified tool integration.

InfoSeeker: Scalable Hierarchical Agent Framework for Web Search

Discover InfoSeeker, a scalable hierarchical agent framework improving web information seeking with faster, accurate parallel processing.

FoE: Why First Solutions Excel in Large Reasoning Models

Discover how the Forest of Errors (FoE) makes first solutions the best in large reasoning models, improving accuracy and efficiency with the RED framework.

AgentHazard: Benchmark for Detecting Harmful Agent Behavior

AgentHazard benchmark evaluates harmful behavior in computer-use agents, highlighting safety risks and the need for improved safeguards in AI models.

Optimality of Large Language Models in AI Planning Tasks

Explore how large language models achieve optimal planning in complex AI tasks, outperforming traditional algorithms with advanced reasoning.

Popular

Subscribe