An AI reviewer can leave dozens of comments and still miss a bug that matters for a release. On October 5, 2026, GitHub introduced ReviewBench, a research preview of a benchmark that compares agents by the issues they find, the ones they miss, and the accuracy of their feedback.
For developers, the important shift is from counting comments to checking their substance. The benchmark compares an agent’s output with a reference set of issues for each pull request, and separately accounts for the severity and category of findings.
What ReviewBench evaluates
The dataset includes 219 public pull requests from 187 openly licensed repositories, with code in 19 languages. GitHub says it analyzed the distribution of more than 103.9 million PRs when assembling the corpus. Repository and language sizes were selected with GitHub’s overall sample in mind, while the size of the changes was intentionally skewed toward more substantive PRs suitable for review.
For each PR, the benchmark stores reference findings labeled by severity and category—for example, correctness, security, reliability, maintainability, and testing. ReviewBench measures precision: what share of an agent’s comments correspond to actual issues; and recall: what share of known issues the agent found. The Fβ score lets users adjust the relative weight of these two criteria: greater emphasis on recall suits finding more issues, while greater emphasis on precision helps reduce noisy feedback.
How to interpret GitHub’s figures
The publication reports 96.6% agreement between ReviewBench annotations and an independent re-review by senior engineers. This measures consistency in expert assessments of reference findings, not the accuracy of any AI agent.
Separately, GitHub described an internal A/B test of a model ensemble for Copilot code review. According to the company, compared with the control group, the share of comments followed by corresponding code changes rose by 8.0%, recall increased by 13.6%, comment volume rose by 61%, and review cost fell by 8.0%. These are results from a single GitHub experiment; they are not an evaluation of all agents or a promise of the same effect for other teams.
Separating the metrics is especially useful when a team is choosing review settings. A high volume of feedback may come with more useful findings, but volume alone does not show how many were confirmed. ReviewBench also breaks down results by severity and category, so teams can examine, for example, critical bugs or security concerns separately.
How to try the benchmark
According to GitHub’s description, the research preview lets users explore the full dataset, compare published results, and run their own agent. A preliminary run uses 25 PRs; a full run covers 219 PRs across three rounds. Participants provide a container image, configuration, and their own model key, while evaluation uses a shared judge. Results remain private until the submission has been reviewed and approved.
The practical value of such a run is to establish a comparable baseline, then test the system against a team’s own review rules and the kinds of changes typical for its repository. A particular team’s conditions—languages, architecture, severity thresholds, and acceptable noise levels—may differ from the makeup of the general corpus.
GitHub says that, in its experiments, changes in ReviewBench offline evaluations aligned directionally with results in production. The published measurements come from the company’s own tests. So for now, the benchmark is most reliably used to compare agents on a shared dataset and diagnose their strengths; decisions about adoption should also be based on testing in the specific team’s workflow.