ConfidenceBench: Benchmarking LLM Confidence Calibration Across Leading Models
A new arXiv preprint introduces ConfidenceBench, a benchmark designed to evaluate how well large language models (LLMs) verbalize their confidence in answers using Brier scores. Testing 15 prominent LLMs, the study finds that top-performing models in accuracy are not always the best-calibrated in confidence, with some models showing severe miscalibration. The benchmark works via prompting, making it applicable to both open- and closed-source models.
Why it matters: This work highlights that confidence calibration is a distinct and critical aspect of LLM reliability, with implications for deploying these models in high-stakes or safety-sensitive applications.
Full story at: arXiv AI/ML ↗