Quantifying Ranking Uncertainty in LLM Benchmarks
A new preprint analyzes sources of uncertainty in the MMLU benchmark and proposes modifications to hypothesis tests for ranking large language models (LLMs). The authors demonstrate that ranking variability across MMLU subjects is substantial and argue that this variability should be considered when comparing LLMs or identifying top-performing models.
Why it matters: This work provides a statistical framework to quantify uncertainty in LLM leaderboard rankings, enabling more rigorous and reliable model comparisons.
Full story at: arXiv Machine Learning ↗