MedMT-Bench: Testing LLMs on Long Medical Conversations

MedMT-Bench evaluates LLMs' ability to handle long multi-turn medical conversations, revealing current AI limitations in clinical dialogue understanding.

Cluster-R1: Instruction-Following Large Reasoning Models

Discover Cluster-R1, a novel approach where large reasoning models follow instructions to improve autonomous data clustering and interpretation.

Symbolic-Mechanistic Evaluation for AI Beyond Accuracy

Discover a symbolic-mechanistic approach to AI evaluation that reveals true model generalization beyond misleading accuracy metrics.

MSA: Efficient Memory Sparse Attention for 100M Token AI Models

Discover MSA, a scalable memory sparse attention model enabling efficient AI processing of up to 100 million tokens with minimal performance loss.

Medical Coding with LLMs Using Privacy-Preserving Synthetic Data

Enhance medical coding accuracy using large language models trained on privacy-preserving synthetic clinical data for safer, efficient healthcare automatio...

Popular

Subscribe