A new preprint examines how a coding agent learns across a sequence of abstract reasoning tasks from the ARC-AGI-3 benchmark. The agent operates with a frozen foundation model inside a fixed harness, meaning its core parameters are not updated during the tasks. Instead, it acts by writing and running Python and shell scripts.

Because the agent retains no state, any learning must emerge from the interaction between the agent, the scripts, and the task environment. The authors trace the agent's 'thoughts' to understand how it improves across tasks, framing their findings as lessons for continual learning—particularly how to build systems that adapt without modifying a frozen model.

As a single preprint, these conclusions are preliminary, but they point toward a promising direction for continual learning without weight updates.