Large language model–based web agents are typically judged by automated evaluators that inspect only the final outcome. A new preprint argues that this narrow focus may not capture whether an agent actually completed a task correctly or merely produced a superficially correct result.

The study, posted on arXiv, conducts an audit of WebArena-Lite, a benchmark for web agent tasks. Instead of relying solely on rule-based or language-model checks, the authors bring in human reviewers to verify task completion and examine the full trajectories agents take. This dual review—outcomes and paths—offers a more granular view of agent performance.

Because the source is a single preprint and the abstract is truncated, full findings are not yet available. Still, the audit underscores a known concern: automated evaluation can diverge from human judgment, especially when intermediate steps matter or when multiple strategies lead to the same final state. The authors' emphasis on trajectory-level review suggests that outcome-only metrics may be insufficient for reliable web agent assessment.