Calibration-Family Overfit: Trusted Sabotage Monitors Fail to Transfer Across Model Lineages
A new preprint demonstrates that trusted sabotage monitors, used to detect harmful actions in AI models, exhibit significant calibration-family overfit: their effectiveness drops sharply when applied to models from a different lineage than the one they were trained on. In a code-backdoor detection task, an off-lineage monitor detected only 19% of attacks compared to 41% for an in-lineage monitor at a 1% audit budget. The authors also propose a four-step protocol to address this transfer gap.
Why it matters: This work exposes a critical limitation in current AI safety evaluation practices, showing that monitor accuracy on a single model pairing can significantly overstate real-world safety across diverse model lineages.
Full story at: arXiv Cryptography and Security ↗