A new paper on arXiv investigates a subtle failure mode for machine learning classifiers: the possibility that their reliability—the relationship between confidence and accuracy—changes even while the overall distribution of confidence scores stays exactly the same. The authors formalize this as a worst-case problem, asking how much the reliability relation can move under covariate shifts that leave the confidence distribution fixed.

The work highlights a gap in common evaluation practices. If a model's confidence distribution appears stable, practitioners might assume its calibration is also stable. But this research shows that conditional accuracy can drift substantially under such conditions, meaning the model could become overconfident or underconfident in ways that are invisible to standard distribution-level checks.

By focusing on the worst-case movement of reliability, the paper provides a theoretical bound on how bad this drift can get. This offers a cautionary note for model monitoring and suggests that maintaining a fixed confidence distribution is not sufficient to guarantee stable reliability in deployment.