A new paper on arXiv argues that looped Transformers—models that reuse the same layers repeatedly at test time—require their own residual connections to avoid performance collapse. The authors observe that simply increasing the number of loop iterations can reduce reasoning accuracy, a counterintuitive result given the usual assumption that more computation helps.
To fix this, they introduce what they call loop-native attention residuals. These are designed specifically for the looped setting, rather than borrowing standard residual paths from non-looped Transformers. The goal is to allow scaling to tens of thousands of test-time iterations without the accuracy degradation seen in naive looping.
The abstract is brief and does not include experimental details or comparisons, so the claims are preliminary. Still, the work highlights a practical obstacle in test-time scaling: more iterations are only useful if the model remains stable across them.