← Back to brief
ResearchOfficialPreprintarXiv Computation and Language

Study Finds LLM Ensemble Diversity Metrics Often Reflect Capability, Not True Diversity

A new arXiv preprint audits five commonly used diversity metrics for selecting large language models (LLMs) in ensemble majority voting. The study finds that these metrics are largely entangled with model capability, rather than measuring true diversity. After controlling for capability, only a modest residual link remains between shared errors and ensemble voting gains, challenging the assumption that diversity metrics reliably predict ensemble improvement.

Why it matters: This result questions the reliability of widely used diversity metrics in LLM ensemble construction, potentially impacting how practitioners combine models for better performance.

Full story at: arXiv Computation and Language