Large language models are increasingly being used as algorithm-design agents that write and refine algorithms. Many of these systems rely on in-context evolution, where the model iteratively improves an algorithm using only the examples and feedback contained in its own context. The new paper argues that this approach, while successful in some settings, can reach a plateau where further generations produce no meaningful progress.

The paper proposes Population-Curated Policy Optimization as a remedy. Rather than relying on pure in-context evolution, the method curates a population of algorithm designs and uses their collective outcomes to update the policy. This is intended to push the agent past the point where in-context evolution typically stalls, allowing self-improvement to continue over longer horizons.