Self-improving LLM systems work by proposing changes to themselves and keeping those that score better on a small evaluation set. A new arXiv paper treats this keep-if-better step as a selection problem under measurement noise, drawing attention to a subtle but important failure mode.
The authors argue that when the evaluation set is small, the observed score advantage of a proposed change may be partly due to random noise rather than genuine improvement. This is the 'winner's curse': the change that looks best is often the one whose score was most inflated by luck. Over repeated self-improvement loops, this can lead to lock-in, where the system commits to a change that is actually worse than the alternative.
The paper models the correlated errors of the evaluation process and examines how acceptance rules—the criteria for keeping a change—can be designed to reduce the risk of selecting noisy winners. The findings suggest that naive keep-if-better strategies are fragile, and that more conservative acceptance rules may be needed for reliable self-improvement.
This work is a cautionary note for the growing field of recursive LLM improvement. It highlights that without careful statistical controls, self-improving systems may not just fail to improve—they may actively degrade by locking in on measurement artifacts.