← Back to brief
ResearchOfficialPreprintarXiv Computation and Language

Systematic Study Finds Universal Response Drift Across Leading LLMs

A systematic human evaluation of ten leading large language models (LLMs) across 62 multidomain questions found that all models exhibit response drift—deviations from expert-validated references. Most models showed high rates of drift (78-81%), while two had notably lower rates (47-49%). The extent and pattern of drift varied by domain and question, and automated metrics accounted for less than 2% of the variance in human judgments.

Why it matters: The findings highlight a widespread and domain-dependent limitation in current LLMs that cannot be reliably detected by automated metrics alone, underscoring the need for human evaluation in assessing model reliability.

Full story at: arXiv Computation and Language