Standard evaluation of scientific-search agents relies on final-answer accuracy, but that single number cannot reveal whether an agent ever encountered the target paper, tried to inspect it, or accepted an answer after looking at it. A new preprint introduces decision-checkpoint auditing, a method that records these intermediate steps rather than only the final output.

The checkpoints separate three failure modes: the agent never saw the relevant paper, saw it but did not attempt to inspect it, or inspected it and still returned an accepted answer. This lets developers see exactly which stage of the pipeline is failing.

Because this is a single preprint, there are no independent results to compare against. The authors position the approach as a diagnostic tool for building more reliable search agents.