Understanding Grokking: From Abstraction to AI Intelligence

Explore how AI models evolve from memorization to intelligence through grokking, driven by structural simplification and complexity measures.

Reliability Science Framework for Long-Horizon LLM Agents

Discover a new reliability science framework to evaluate long-horizon LLM agents beyond pass@1, improving consistency across extended tasks.

Xuanwu VL-2B: Industrial-Grade Multimodal AI for Content

Discover Xuanwu VL-2B, an industrial-grade multimodal AI model enhancing content moderation with fine-grained perception and robust adversarial handling.

RIDE Study: Impact of Routing Meta Prompts on LLM Stability

Explore how routing-style meta prompts affect density, stability, and performance in large language models, challenging traditional routing beliefs.

AEC-Bench: AI Benchmark for Architecture & Construction

Discover AEC-Bench, a multimodal AI benchmark for evaluating agentic systems in architecture, engineering, and construction projects.

Popular

Subscribe