GitHub has released ReviewBench, a research-preview benchmark that quantifies how effectively AI code review agents spot problems before software ships. The benchmark runs each agent against the same 219 pull requests drawn from 187 open-source repositories, deliberately weighting toward multi-file changes where review quality matters most. Users can inspect the test data, reproduce results, and filter leaderboard rankings by issue severity, category, and scoring preference.
The core evaluation relies on a curated "golden set" of validated findings, built from human reviewers, follow-up code changes, analysis tools, and AI models. ReviewBench reports grounded precision and recall—how many of an agent's findings match known issues and how many known issues it detects—plus augmented versions that give credit for valid findings outside the reference set. An AI judge assesses those extras, so reviewers are not penalized for catching problems the golden set missed.
In a production experiment with Copilot code review's lite tier, GitHub combined several independent model runs into a single review. That change improved benchmark scores and real-world results: the share of comments judged to prompt code changes rose 8%, recall climbed 13.6%, and cost per review fell 8%. Feedback also shifted toward critical and moderate issues, with fewer minor suggestions. GitHub stresses that user testing remains the final measure of impact, but the benchmark helps decide which updates are worth that testing.
Researchers and practitioners can submit their own agents by registering a container image and configuration. A 25-pull-request test set allows quick iteration before running the full 219-PR suite. Scores stay private until a maintainer approves the submission, then appear on the public leaderboard. The benchmark's authors invite scrutiny of its assumptions and methodology, signaling that ReviewBench is meant to evolve alongside the tools it measures.