← Back to brief
ResearchOfficialPreprintarXiv Computers and Society

New Benchmark Finds LLMs Underperform on Women's Health Scenarios, Top Model Scores 72.1%

A new arXiv preprint introduces WHBench, a benchmark of 47 expert-designed scenarios covering 10 women's health topics, to evaluate large language models (LLMs). Testing 22 models, researchers found that none surpassed 75% mean performance, with the best model achieving 72.1%, and identified clinically relevant failure modes such as outdated advice and safety issues.

Why it matters: The results highlight significant safety and reliability gaps in current LLMs for women's health, emphasizing the need for expert oversight before clinical use.

Full story at: arXiv Computers and Society