← Back to brief
ResearchOfficialPreprintarXiv AI/ML

Wisdom of LLM Crowds: Aggregation and Contamination in Language Model Ensembles

Researchers investigated whether aggregating probability estimates from multiple large language models (LLMs) can outperform individual models. They found that learned aggregation methods, such as logistic regression and multilayer perceptrons, consistently outperformed both individual models and classical aggregation techniques. However, the study also revealed that training data contamination significantly inflated the apparent performance gap between frontier and smaller models, which shrank from 35.8% to 8.9% when using uncontaminated data.

Why it matters: This work demonstrates both the potential and the pitfalls of ensemble approaches for LLMs, emphasizing the importance of contamination-free evaluation for accurately measuring model capabilities.

Full story at: arXiv AI/ML