· via GitHub Blog
GitHub launches ReviewBench, an open benchmark for AI code review agents
GitHub has published ReviewBench, an open benchmark that scores AI code review agents on 219 real-world pull requests using grounded and augmented precision-recall metrics.

GitHub opens a public yardstick for AI code review agents
GitHub has launched ReviewBench, an open benchmark for scoring AI code review agents against realistic pull requests. Writing on the GitHub Blog, the company frames the release as an answer to a persistent problem: existing evaluations tend to force tradeoffs between label quality, coverage and how faithfully they represent day-to-day review work, leaving teams without a dependable way to compare reviewers or to tell whether a change to their own agent will actually help in production.
A corpus modeled on 103.9 million pull requests
According to GitHub, the benchmark's shape is grounded in an analysis of 103.9 million real pull requests on the platform, used to map how review workloads distribute across language, repository size and change size. The published corpus itself contains 219 pull requests from 187 public repositories under open-source licenses, spanning 19 languages, with language and repository-size profiles aligned to GitHub-wide figures.
One distribution is intentionally shifted. Pull request size is weighted toward substantive, multi-file changes and away from tiny single-file edits, which GitHub says are overrepresented in the raw data and matter least for measuring review quality. The complete dataset is available publicly.
Ground truth from many hands
Because no single reviewer, human or model, catches everything, GitHub builds its golden set in three stages. Candidate findings are first collected from human reviewers, from issues implied by authors' follow-up commits, from deterministic analysis tools, and from several frontier LLMs across model families. Semantically duplicate findings are then merged, so overlap between sources cannot artificially inflate the set. Finally, every candidate is judged against one shared rubric: it counts as a true positive only if it is genuine, relevant and substantive, no matter where it came from.
Claude Sonnet 5 acts as the grader applying that rubric, and GitHub publishes both the rubric and the judge for transparency and reproducibility. Senior engineers independently labeled golden true positives before release, with GitHub reporting 96.6% agreement, and a human-labeled development set is used to calibrate the automated grader against human judgment.
Grounded and augmented metrics
ReviewBench reports two families of metrics. Grounded precision, recall and F1 compare an agent strictly against the existing gold labels: of the issues already known, how many did the agent find, and how much of what it reported matched. Augmented precision, recall and F1 go further by sending each unmatched finding to the judge, which decides whether it is a genuine discovery or a false alarm. That distinction matters because any fixed golden set becomes stale as agents find issues its creators never anticipated, and augmented scoring gives credit for those findings rather than punishing them automatically.
Because augmented recall expands its denominator with whatever each agent discovers, GitHub designates grounded recall as the headline figure for cross-system comparison, with augmented numbers treated as per-system diagnostics.
Results can be sliced by severity (critical, medium and low) and by category, including correctness, security, reliability, maintainability and testing. The F-beta weighting is also adjustable, letting evaluators favor either broad coverage or precision. GitHub argues this configurability is necessary because there is no single ideal review profile: some developers want only critical problems surfaced, while others accept more noise in exchange for wider coverage.
GitHub adds that ReviewBench movement has tracked the direction of its online production experiments with Copilot code review, which is what gives the company confidence that offline gains translate into real user benefit. Teams building their own review agents can onboard them and submit results.
Why it matters
For anyone buying or building AI review tooling, ReviewBench offers something the field has lacked: a reproducible, public offline signal tied to production behavior rather than a hand-picked demo set. The severity and category slices let organizations match a reviewer to their actual risk tolerance, so security-focused teams and noise-averse teams can compare agents on the numbers they care about.
The design choices are notable in themselves. Multi-source ground truth with deduplication, a published rubric and grader, published human-agreement rates, and augmented metrics that credit novel findings all address familiar criticisms of LLM benchmarks: stale golden sets and opaque judging. Two caveats are worth holding onto. The 219-pull-request corpus is small relative to the 103.9 million requests analyzed for distribution, and the judge is itself an LLM, so scores ultimately rest on grader quality — a risk mitigated, though not eliminated, by calibration and expert audits.
- #ai
- #code-review
- #benchmark
- #github
- #open-source