A new preprint on arXiv describes how autonomous coding agents work: they read code, run commands, edit files, and submit patches to resolve repository issues. The study investigates whether giving these agents more inference-time compute reliably improves their success.
The abstract states that extra inference-time compute yields gains only when it produces a useful repair and supplies reliable... The text cuts off there, so the exact nature of the reliability condition is not available in the abstract. The clear implication is that more compute is not automatically better; it must lead to a genuine fix.
This matters for developers who scale up compute for coding agents. If the extra compute does not produce a useful repair, it may not improve outcomes. The abstract offers a caution against assuming that simply increasing inference-time compute will make agents more reliable.