MAC-Attention: Fast, Accurate Attention for Long-Context LLMs

Discover MAC-Attention, a novel scheme boosting large language models with fast, accurate long-context attention and up to 14x speed improvements.

Diversity-Aware Reverse KL Divergence for LLM Distillation

Explore Diversity-aware Reverse KL Divergence (DRKL) to improve large language model distillation with better performance and output diversity.

QUEST: Robust Query-Modulated Spherical Attention in Transformers

Discover QUEST, a stable and robust query-modulated spherical attention method enhancing Transformer training and performance across domains.

AI Agents: Adoption, Architectures & Insights from Experts

Explore AI agents' adoption, architectures, and expert insights to enhance your AI strategy with real-world practitioner takeaways.

Explainable AI for Blind Users: Trust & Accessibility

Explore how explainable AI can improve trust and accessibility for blind and low-vision users through multimodal, blame-aware designs.

Popular

Subscribe