Theoretical Limits and Certified Fixes for Self-Poisoning in Adaptive Out-of-Distribution Detection
A new arXiv preprint provides a theoretical analysis showing that adaptive out-of-distribution (OOD) detectors, which update from unlabelled data streams, can collapse due to a self-poisoning feedback loop when a key parameter exceeds a threshold. The authors introduce a certified admission gate that provably prevents this collapse, even under adversarial contamination, and present a calibration method for static drift. They also prove an impossibility result: without labels, it is fundamentally impossible to distinguish between drift and contamination, setting a theoretical ceiling for such detectors.
Why it matters: This work exposes a critical vulnerability in a widely used class of OOD detectors and offers certified solutions, with implications for the reliability of AI systems in safety-critical applications.
Full story at: arXiv Machine Learning ↗