Self-Poisoning in Adaptive Out-of-Distribution Detection: A Sharp-Threshold Theory and Certified Label-Free Calibration
Test-time adaptive OOD detectors update memory banks from unlabelled streams, obeying a provable dynamical law modeled as a generalized Pólya urn. The bank impurity converges to a mean-field equilibrium with a slope acting as a reproduction number; below one, impurity is benign, above one, the bank is poisoned and the detector collapses. The admission kernel is affine with a slope just below one, making the detector class near-critical by design, and the predicted threshold matches empirical col
Analysis
TL;DR
- Test-time adaptive OOD detectors update memory banks from unlabelled streams, obeying a provable dynamical law modeled as a generalized Pólya urn.
- The bank impurity converges to a mean-field equilibrium with a slope acting as a reproduction number; below one, impurity is benign, above one, the bank is poisoned and the detector collapses.
- The admission kernel is affine with a slope just below one, making the detector class near-critical by design, and the predicted threshold matches empirical collapse across 96 settings.
- A certified admission gate reading only a frozen reserve severs the feedback loop, removing the transition at every contamination rate while controlling false positives label-free.
- CDC restores nominal FPR label-free under drift, and a two-world impossibility theorem shows drift and contamination are indistinguishable without labels, forcing a closed-form power ceiling.
Why It Matters
This paper provides a theoretical foundation for understanding and mitigating self-pooping in adaptive OOD detection, which is crucial for maintaining the reliability of AI systems in dynamic environments. The findings offer actionable insights for designing robust OOD detectors that can handle both contamination and drift without compromising performance.
Technical Details
- The adaptation of memory banks in test-time adaptive OOD detectors is modeled as a generalized Pólya urn, proving almost-sure convergence to a mean-field equilibrium.
- The slope of the mean-field equilibrium acts as a reproduction number, determining whether impurity stays benign or leads to detector collapse.
- The admission kernel is affine with a slope just below one, indicating that the detector class is near-critical by design.
- A certified admission gate reading only a frozen reserve severs the feedback loop, ensuring the detector remains robust even under adversarial contamination.
- CDC (Certified Drift Calibration) restores nominal FPR label-free under drift, addressing the complementary issue of static-calibration failure.
- The two-world impossibility theorem highlights the indistinguishability of drift and contamination without labels, setting a closed-form power ceiling for label-free adaptive OOD detection.
Industry Insight
- AI practitioners should consider implementing certified admission gates to prevent self-pooping in adaptive OOD detectors, ensuring long-term reliability and performance.
- The near-critical nature of current detector designs suggests that small adjustments in the admission kernel slope can significantly impact detector robustness, warranting careful tuning.
- The theoretical framework provided in this paper can guide the development of more robust OOD detection algorithms, particularly in applications where data streams are unlabelled and subject to drift or contamination.
Disclaimer: The above content is generated by AI and is for reference only.