← Back to brief
Policy & SafetyOfficialPreprintarXiv Cryptography and Security

Calibration-Family Overfit: Trusted Sabotage Monitors Fail to Transfer Across Model Lineages

A new preprint demonstrates that trusted sabotage monitors, used to detect harmful actions in AI models, exhibit significant calibration-family overfit: their effectiveness drops sharply when applied to models from a different lineage than the one they were trained on. In a code-backdoor detection task, an off-lineage monitor detected only 19% of attacks compared to 41% for an in-lineage monitor at a 1% audit budget. The authors also propose a four-step protocol to address this transfer gap.

Why it matters: This work exposes a critical limitation in current AI safety evaluation practices, showing that monitor accuracy on a single model pairing can significantly overstate real-world safety across diverse model lineages.

Full story at: arXiv Cryptography and Security