AIToday
AI Coding AssistantsOpen-Source AIGitHub Blog (AI)Published: Oct 6, 2026, 01:00 JST

GitHub launches ReviewBench, an open AI code review benchmark

GitHub launches ReviewBench, an open AI code review benchmark

3 Key Points

  1. What happened

    GitHub released ReviewBench, built on 103.9 million GitHub pull requests and 219 pull requests from 187 repos in 19 languages, with findings from human reviewers, frontier LLMs, and static analysis labeled by severity and category.

  2. Why it matters

    It gives teams an offline signal that points the same direction as later production experiments, so measured improvements are more likely to reflect meaningful gains for users, hedged as indicative of direction rather than a guarantee.

  3. What to watch

    Whether the benchmark's predictive accuracy holds as agents grow more capable, since its golden set is fixed and can become incomplete. Watch the published leaderboard and the 25-PR test set for new submissions.

WHO IT HITSThis lands on engineering teams that build or buy AI code review agents, giving them a public yardstick to compare systems before committing to production deployments. It also affects senior engineers and maintainers who will now submit their own agents to the leaderboard.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Agentic code review is increasingly used to inspect pull requests and catch issues before code ships, but teams have had a hard time measuring how well these AI reviewers actually work. Existing benchmarks often trade off label quality, coverage, and real-world representativeness, so GitHub built ReviewBench to combine all three. The benchmark analyzes 103.9 million GitHub pull requests to model language, repository size, and even pull request size, deliberately upweighting the reviewable middle and tail so tiny, single-file changes do not dominate. Ground truth comes from a multi-source golden set—human reviewers, follow-up commits, static analysis tools, and frontier LLMs—with duplicate findings merged and every finding validated under a published rubric by an LLM judge. Senior engineers who had not built the dataset re-labeled every ground-truth finding and agreed with ReviewBench 96.6% of the time.

The benchmark reports six metrics in two families: grounded precision, recall, and F1 against the fixed golden set, and augmented precision, recall, and F1 that give agents credit for valid issues not previously known. That matters because a fixed golden set inevitably becomes incomplete as systems get more capable, and augmented metrics avoid penalizing agents for finding unanticipated problems. Users can slice results by severity and category and adjust a β parameter to favor broader coverage or lower noise, with the leaderboard re-ranking accordingly.

GitHub used ReviewBench to evaluate Copilot code review across iterations, and one recent experiment—a multi-model ensemble that combines several independent model runs into a single review—illustrated the benchmark's value. ReviewBench predicted higher precision, recall, and comment volume with lower cost per review; an online A/B test later showed addressed rate rose 8.0%, recall rose 13.6%, comment volume rose 61%, and cost per review fell 8.0%. Severity-level evaluation also matched, predicting a 227% increase in critical comments versus 262% online. The open question is whether this predictive accuracy will hold as review agents evolve and the fixed golden set ages; the test is whether future submissions still see offline signals track production.

FAQ
How can teams use ReviewBench results to choose a code review tool?
ReviewBench lets results be sliced by severity and category, and users can adjust the β in the Fβ score to weigh recall or precision. The leaderboard re-ranks accordingly, helping users identify systems that best match their review priorities.
What did GitHub's A/B test show about the multi-model ensemble review's actual production impact?
Addressed rate rose 8.0%, recall rose 13.6%, and comment volume rose 61%, while cost per review fell 8.0%, all relative to the production control. ReviewBench had predicted the direction of these changes.
How can someone submit their own code review agent to ReviewBench?
Sign in with GitHub on the ReviewBench website, register your agent with a container image and configuration, run against the 25-PR test set, then do a full run of 219 pull requests. Scores are published only if they outperform the agent's current leaderboard score or it's the first entry.
GitHub Blog (AI)Read Original Article

Also reported by GitHub Copilot Blog

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleNorway drafts limits on smart glasses filming