Financial LLM agents are typically evaluated by comparing their end-to-end returns with a baseline and testing the paired difference against zero. According to a new arXiv preprint, this measures only whether deploying the agent changes realized performance, leaving the internal drivers of that change unexamined.
The paper proposes an "agent policy-value audit" that separates transition composition from event selection. This distinction is intended to provide a more detailed view than end-to-end returns alone, allowing evaluators to ask not just whether an agent performed differently, but which part of its decision process contributed to the outcome.
The preprint positions this as a complement to existing evaluation methods. While the abstract is brief, the framing suggests a shift toward diagnostic evaluation of financial LLM agents rather than relying solely on aggregate performance metrics.