Latency-Response Theory Model: Evaluating LLMs via Accuracy and Chain-of-Thought Length
Researchers introduce the Latency-Response Theory (LaRT) model, which jointly models large language model (LLM) response accuracy and chain-of-thought (CoT) length for evaluation purposes. The model incorporates a correlation parameter between latent ability and latent speed, and is shown through theoretical analysis, simulations, and real LLM benchmark data to outperform traditional Item Response Theory (IRT) in estimation accuracy and evaluation efficiency. LaRT also produces different LLM rankings and demonstrates improved predictive power and ranking validity compared to IRT.
Why it matters: This approach could lead to more nuanced and statistically robust assessments of LLM reasoning by leveraging both accuracy and reasoning process length.
Full story at: arXiv Statistical ML ↗