Most evaluations of LLM agents measure whether a task was completed or whether the agent's plan matches a reference plan. According to a new arXiv paper (2608.04265v2), this is insufficient for strategic cyber-physical systems, where the environment includes autonomous participants and real physical processes.
The authors argue that an architecture must remain appropriate after those participants respond and physics takes over. In other words, a plan that looks good on paper may fail once other agents react strategically or the system's dynamics play out. The paper appears to call for evaluation strategies that incorporate these post-decision effects.
Because the abstract is truncated, the specific proposed evaluation method is not described in the available text. The key claim, however, is clear: task success and plan agreement are not enough for these settings.