A new arXiv preprint examines whether benchmarks remain trustworthy when developers use evaluation feedback to iteratively refine their models. The authors focus on settings where model selection happens adaptively across multiple criteria, rather than through a single static comparison.

The study finds that the worst-case test-set size needed to estimate the best score among k adaptively chosen models grows with k. This means that as developers try more model variants in response to feedback, the benchmark's ability to reliably identify the true best performer degrades unless the test set is substantially enlarged.

Because the abstract is truncated, the exact scaling relationship is not fully detailed, but the core implication is clear: benchmarks designed for one-shot evaluation may not support the iterative, feedback-driven workflows common in modern model development. The authors caution that without accounting for this adaptive selection, reported scores can be misleading.

This work highlights a practical limitation of current evaluation practices, though it does not propose a specific remedy. As the only source, it offers a focused warning rather than a comparative analysis.