An arXiv preprint asks how much weight a coding agent's self-correction should carry when the agent learns from its own revisions. The authors propose an adaptive Dirichlet evidence framework for self-distillation, in which each correction's learning weight is tied to the evidence behind it.

Execution feedback lets coding agents revise programs and learn from those corrections. The paper's key move is to make the weight of a correction depend on both the transitions supported by its executions and the amount of evidence supporting those transitions. In other words, not all self-corrections are equally trustworthy.

By tying learning weight to evidence, the approach aims to make self-correction a more reliable training signal. The preprint is available on arXiv under identifier 2610.08514v1.