Every self-improving agent loop faces a recurring question: did the change actually make the agent better? According to a new arXiv paper, the answer always comes from a verifier. But for open-ended tasks, no such verifier exists, so the loop is handed a hand-written rubric instead.

The paper's title — "Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents" — points to a proposed solution: let the graders evolve together with the agents, while keeping them inspectable. This directly addresses the blind spot in current self-improvement loops, where the rubric itself is never questioned.

The abstract is brief, and the full method is not described in the available text. Still, the framing is significant: it highlights that the verifier, not just the agent, may be the limiting factor in open-ended self-improvement.