WHBench: Benchmarking LLMs for Women’s Health AI Safety

WHBench evaluates large language models on women's health topics using expert validation to ensure clinical accuracy, safety, and equity in AI medical guid...

LLM Evaluation Validity for Business in Conversational Commerce

Explore how LLM-based dialogue evaluation correlates with business outcomes in conversational commerce, improving AI-driven sales strategies.

How Language Models Process Ethical Instructions: Key Insights

Explore how top language models process ethical instructions, revealing distinct types of ethical reasoning and the impact of instruction formats.

RiDiC Dataset for Long-Form Factuality Evaluation

Explore the RiDiC dataset designed to evaluate long-form factuality in multilingual LLMs using controlled popularity distributions.

Entropy-Guided Decoding to Boost LLM Reasoning Accuracy

Enhance large language model reasoning with entropy-guided decoding, improving accuracy and efficiency over traditional methods like greedy decoding.

Popular

Subscribe