Study Finds Temperature Sampling in LLMs Yields Limited Uncertainty Compared to Model Ensembles
A new arXiv preprint reports that repeatedly sampling answers from a single large language model (LLM) at high temperature produces only one dimension of meaningful variation, while using an ensemble of 24 different models reveals four. The analysis, conducted across several benchmarks, suggests that temperature-based sampling provides per-question uncertainty but lacks the richer, cross-question uncertainty structure captured by diverse model ensembles.
Why it matters: This challenges the common practice of using temperature-based sampling for uncertainty estimation in LLMs, indicating it cannot substitute for the broader epistemic coverage provided by model ensembles.
Full story at: arXiv AI/ML ↗