BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges
BACON is a four-stage pipeline that integrates budgeted human calibration with outputs from multiple AI judges to generate more accurate and calibrated annotations. The method constructs auxiliary features from AI judges, collects human labels on a small subset, and trains a model to produce calibrated predictions. BACON demonstrates improved predictive accuracy, ranking consistency, and reduced bias and variance compared to using raw AI outputs or only human labels.
Why it matters: This approach offers a scalable and statistically robust framework for evaluation tasks where human annotation is limited, addressing biases in AI judge outputs that can affect model assessment and reporting.
Full story at: arXiv Machine Learning ↗