Large reasoning models (LRMs) are designed to tackle complex tasks by thinking longer and more carefully. A new preprint from arXiv shows that this strength can become a weakness: the very process that powers their advanced reasoning can spiral into uncontrolled loops.

The paper describes how LRMs sometimes get stuck in redundant verification and persistent generation cycles. Instead of reaching a conclusion, the model keeps re-checking its own steps or producing repetitive outputs, which inflates the computational cost without improving the final answer.

The findings suggest that while extended reasoning is generally beneficial, it needs guardrails. Without them, models may waste resources and produce less reliable results—an important consideration for deploying LRMs in real-world applications.