Therapeutic antibody and TCR discovery increasingly relies on AI models that predict binding, but these models often perform unpredictably when applied to new datasets, leading to low discovery rates. This uncertainty makes it hard to know which models to trust and when to retrain them.

A new approach called CaliPPer aims to address this by quantifying, predicting, and improving model performance for binding prediction. The method appears to leverage density-ratio techniques, similar to existing label-free estimators such as PAPE and M-CBPE, to assess how well a model will generalize without requiring ground-truth labels on the target dataset.

While the full details are not yet available, the core idea is to give researchers a practical tool for anticipating model failures and guiding model selection or improvement. This could reduce wasted effort in discovery pipelines and make AI-driven binder design more reliable.