Large language models are increasingly used as autonomous agents that operate over long, multi-turn interactions. In such settings, they must keep track of original objectives, decide when to use tools, and respond to other agents—all while avoiding distractions. The paper argues that existing evaluation methods do not adequately capture these demands, particularly the ability to delay immediate gratification in favor of long-term success.
To fill this gap, the authors propose a multi-agent survival micro-benchmark. The benchmark is designed to test whether an LLM agent can maintain its goals over extended interactions when exposed to social pressures, assigned personas, and limited tool budgets. These factors are treated as controlled variables, allowing researchers to isolate how each affects an agent's capacity for sustained, goal-directed behavior.
The work is positioned as a step toward more auditable evaluation of long-horizon LLM agents. Because the abstract is truncated, the full methodology and results are not available in the source, but the stated aim is clear: create a reproducible test for a capability that is often assumed but rarely measured directly. The paper does not compare against other benchmarks, so its claims stand on their own as a proposal rather than a demonstrated improvement.