A new preprint argues that reviewing an AI-generated study's final artifacts is not enough. Some defects can be spotted in the outputs themselves, but others only become clear when you know what was approved before the study ran. The authors call this the output-only review problem.

Their proposed solution is a 'study contract' — a binding record of declared experimental choices, run obligations, and other pre-execution commitments. By tying the agent to what was originally specified, reviewers can check not just whether the output looks plausible, but whether the execution matched the plan.

The paper is careful to distinguish what output-only review can and cannot verify. The abstract notes that some defects are identifiable from artifacts, while others require knowledge of the pre-execution approval. The study contract is designed to close that second gap, though the full details of how such contracts would be enforced are not available in the abstract.