HealthAdminBench: Benchmarking AI in Healthcare Admin Tasks

Discover HealthAdminBench, a benchmark evaluating AI agents on healthcare administration tasks like prior authorization and appeals management.

GLEaN: Visual Bias Detection in Text-to-Image Models

Discover GLEaN, a scalable approach that visually reveals biases in text-to-image AI models for easy public understanding and transparency.

EE-MCP: Self-Evolving GUI Agents with Automated Learning

Discover EE-MCP, a self-evolving agent framework using automated environment generation and experience learning to improve GUI and MCP task automation.

AI-Driven In-Situ Defect Detection in Wire-Arc AM

Discover how agentic AI enhances in-situ defect detection in wire-arc additive manufacturing for improved accuracy and real-time monitoring.

What Logits Reveal About AI Models: Surprising Insights

Discover how logits in AI models can unintentionally leak sensitive info and what it means for privacy, security, and ethical AI development.

Popular

Subscribe