A new paper on arXiv draws attention to a blind spot in how the field measures large language models. As LLMs are increasingly deployed as agents, they rely on harnesses that handle bounded context, persistent memory, tool use, verification, and repeated execution. These harnesses consume real computational resources, yet the paper argues that existing notions of model capability do not quantify any of it.
The authors suggest this gap could distort progress. A model that scores well on conventional benchmarks may still be impractical if its agentic harness requires excessive compute. The paper calls for evaluation frameworks that account for the full resource footprint of agentic behavior, rather than treating capability and cost as separate concerns.