AIToday
Large Language ModelsAI Coding AssistantsOpen-Source AIGitHub Copilot BlogPublished: Oct 6, 2026, 01:00 JST

GitHub opens ReviewBench, a 219-PR code review benchmark

GitHub opens ReviewBench, a 219-PR code review benchmark

3 Key Points

  1. What happened

    GitHub released ReviewBench, a benchmark of 219 public pull requests across 19 languages, built by analyzing 103.9 million GitHub pull requests and validated by senior engineers who agreed with it 96.6% of the time.

  2. Why it matters

    ReviewBench scores agents on both known and newly discovered issues, and GitHub says its offline results have consistently pointed in the same direction as later production experiments.

  3. What to watch

    The test is whether other teams' agents show the same offline-to-online alignment. Watch the 25-PR test set, then the full 219-pull-request run scored by the published judge.

WHO IT HITSEngineering leaders and developer-tooling teams choosing or building AI code review agents get a shared yardstick with severity and category breakdowns, so they can compare reviewers against their own precision-versus-coverage preferences.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Existing benchmarks for AI code reviewers tend to trade off label quality, coverage and real-world representativeness, which leaves teams without a reproducible way to compare systems. ReviewBench is GitHub's attempt to close that gap by modeling its corpus on over 100 million real pull requests on GitHub and weighting pull-request size toward the reviewable middle and tail rather than tiny single-file changes.

Its ground truth is assembled from human reviewers, issues inferred from author follow-up commits, deterministic analysis tools and multiple frontier LLMs, then deduplicated and judged under one published rubric. Reviewer agents are scored on two metric families: grounded precision and recall against existing gold-set labels, and augmented precision and recall that can credit valid findings no gold-set producer surfaced. Results can be sliced by severity and category, or re-ranked by adjusting β in the Fβ score for teams that prefer broader coverage or less noise.

GitHub says it has used the benchmark to evaluate Copilot code review across successive iterations, and that offline changes have consistently pointed in the same direction as later production A/B tests. The stakes therefore hinge on whether that offline-to-online alignment holds for agents built outside GitHub, which the public leaderboard and self-serve runner are designed to test.

FAQ
How was ReviewBench's ground truth validated?
Senior engineers who had not built the dataset independently re-labeled every ground-truth finding from scratch. Their true/false-positive judgments agreed with ReviewBench 96.6% of the time.
How can a team submit its own code review agent?
Sign in with GitHub on the ReviewBench website and register the agent with a container image, configuration and its own model key. After trying the 25-PR test set, a final run covers all 219 pull requests in three rounds, scored by the same judge.
What did GitHub's own A/B test of an ensemble review show?
Addressed rate rose 8.0%, recall rose 13.6% and comment volume rose 61%, while cost per review fell 8.0%, all relative to the production control. Severity-level evaluation predicted a 227% increase in critical comments, against 262% online.
GitHub Copilot BlogRead Original Article

Also reported by GitHub Blog (AI)

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleSim-trained model with 90.2% success fails AWSIM, 0/4 finish