Skip to content
worth noting Coding

GitHub has released ReviewBench for evaluating AI code review

only one source so far

The ReviewBench benchmark is available in research preview for comparing AI systems for code review and evaluating your own agent. The public dataset contains 219 pull requests from 187 repositories and covers 19 programming languages.

GitHub has released ReviewBench, a benchmark for evaluating the quality of AI systems for code review. In its research preview phase, it allows users to explore a public dataset, compare agents and submit their own system for evaluation. The dataset contains 219 pull requests from 187 public repositories with open-source licenses and covers 19 programming languages.

According to GitHub, the selection is based on an analysis of 103.9 million pull requests. The distribution of languages and repository sizes is intended to reflect overall activity on the GitHub platform. In terms of change size, however, the benchmark deliberately limits the predominance of small edits and gives more weight to larger changes spanning multiple files. Results can be filtered by the severity and category of findings, and the balance between detecting issues and limiting false positives can be adjusted.

According to GitHub, an internal audit by senior engineers who were not involved in creating the dataset agreed with the original evaluation in 96.6 % of cases. Versions of both the dataset and components of the evaluation procedure are tracked to enable repeatable comparisons. The company says that, when developing the Copilot code review feature, the benchmark helps it track improvements, detect declines in quality and predict the direction of results from experiments in production; this reflects the experience of the company with its own product.

What changed

Why it matters

System rankings may change depending on whether a user prioritizes serious errors, broader coverage or fewer false positives. ReviewBench allows these preferences to be reflected in the comparison. Its results should be interpreted with its deliberately higher representation of larger code changes in mind.

Two audiences, two different impacts

What this means

01

For individuals

When choosing an AI assistant for code review, a developer can compare systems based on the severity of findings and their tolerance for false positives.

What to do On the ReviewBench website, compare results for the severity levels and issue categories you need to detect during code review.
More practical updates →
02

For a business

Teams developing their own code review agents gain a public resource for repeatable evaluation of individual versions and tracking declines in quality before deployment.

Development
What to decide Evaluate the current and modified versions of your own agent using the same versions of the dataset and evaluation procedure.
More business impacts →
Copilot Code Review GitHub ReviewBench

Check the original

Event sources

only one source so far · 1 publisher, 0 independent. We count feeds from the same owner only once.

1
GitHub Blog primary source · first detected ReviewBench: An open benchmark for AI code review