A new preprint on arXiv (2610.04168) proposes operational criteria for evaluating agentic large language model (LLM) systems. The authors note that such systems are commonly implemented as an LLM in a loop with planning, memory, tools, and control flow, and they offer "agentic cognitive depth" as a way to assess these designs.
The abstract frames this as an application-focused view that connects agentic LLM research with deployable systems, though the full details of the criteria are not included in the summary. The significance is in shifting evaluation from task-specific outputs toward structural properties of how agents are built.
Because only the abstract is available, the precise operational measures remain unclear, and no comparison with other evaluation frameworks is provided in the source.