Repeated evaluation is often used to estimate benchmark scores, but an accurate point estimate does not by itself guarantee that the reported uncertainty is honest. A new paper on arXiv formalizes this gap, showing that while repeated runs can converge to the right score, certifying a narrow uncertainty interval is a stricter requirement.
The authors characterize this requirement on a fixed grid of M tasks, where each task has L binary paths. Under a hard budget on the number of evaluations, they identify sharp limits on when honest uncertainty can be claimed. The key finding is that replication—running the same tasks multiple times—is necessary to certify narrow uncertainty, even when the aggregate score looks stable.
The work highlights a practical distinction between estimation and certification in benchmark evaluation. For researchers who want to report not just a score but a defensible confidence interval, the results suggest that replication is not optional but a structural requirement. The abstract is truncated, so the full characterization is not yet visible, but the stated direction is clear: honest uncertainty has a cost that repeated evaluation alone cannot avoid.