← Back to brief
Policy & SafetyOfficialPreprintarXiv AI/ML

Rater State Bias in RLHF Preference Data: An Audit Framework

A new preprint identifies a structured confound in Reinforcement Learning from Human Feedback (RLHF): the emotional or psychological state of human raters during annotation can systematically bias preference labels, especially under stressful conditions. These state-dependent shifts are distinct from random noise and can propagate through reward modeling and policy optimization. The authors propose an audit framework with falsifiable predictions and outline a pilot study plan to detect such bias in publicly available instruction-tuned models.

Why it matters: This work draws attention to a novel, testable source of bias in RLHF that could undermine the reliability of preference-based AI alignment methods.

Full story at: arXiv AI/ML