A new arXiv paper, EvalResearchBench, tackles a bottleneck in recursive self-improvement (RSI): the cost of evaluation. RSI systems need feedback to measure progress, but repeatedly running complex benchmarks is expensive and slows down iteration.

The abstract notes that human experts reduce this cost, though the sentence is cut off in the announcement. The paper's title and framing suggest the goal is to let AI agents design their own evaluations, potentially removing a human bottleneck.

Because the abstract is truncated, the full methodology and results are not available from this source. The paper appears to be a benchmark for evaluating agents' ability to generate evaluation tasks, which would be a step toward cheaper, faster self-improvement.