A new paper posted on arXiv examines what it takes for AI agents to stay reliable over long-horizon workflows. The authors argue that such tasks force models to repeatedly act based on changing state, while the context window expands, sub-task difficulty shifts, and new information arrives. Each of these factors, they suggest, represents an independent axis along which agent performance can degrade.

The paper frames these axes as a foundation for testing, rather than assuming a single monolithic failure mode. By separating out context growth, complexity changes, and data arrival, the work offers a way to diagnose where long-horizon agents actually break down. As an arXiv preprint, it has not yet been peer-reviewed, but it points toward a more structured approach to evaluating reliability in agentic systems.