Olfactory Perception Benchmark for Large Language Models

Discover the Olfactory Perception benchmark to evaluate large language models' ability to reason about smell across multiple tasks and languages.

Optimizer-Aware Online Data Selection for LLM Fine-Tuning

Discover a two-stage optimizer-aware method for online data selection that boosts large language model fine-tuning efficiency and performance.

AI-Driven Particle Physics Analysis with LEP Open Data

Discover how AI agents collaborate with physicists to analyze LEP open data, enhancing experimental particle physics measurements and research.

HippoCamp: Benchmarking AI Agents for PC File Management

Discover HippoCamp, a benchmark evaluating AI agents' contextual reasoning and multimodal file management on personal computers with real-world data.

Do Language Models Think Before Deciding? New Insights

Discover how early-encoded decisions shape reasoning in language models, revealing new insights into AI decision-making and cognitive processes.

Popular

Subscribe