
What happened
GitHub released ReviewBench, a benchmark of 219 public pull requests across 19 languages, built by analyzing 103.9 million GitHub pull requests and validated by senior engineers who agreed with it 96.6% of the time.
Why it matters
ReviewBench scores agents on both known and newly discovered issues, and GitHub says its offline results have consistently pointed in the same direction as later production experiments.
What to watch
The test is whether other teams' agents show the same offline-to-online alignment. Watch the 25-PR test set, then the full 219-pull-request run scored by the published judge.
WHO IT HITSEngineering leaders and developer-tooling teams choosing or building AI code review agents get a shared yardstick with severity and category breakdowns, so they can compare reviewers against their own precision-versus-coverage preferences.
Summaries like this, in your inbox every morning.
Existing benchmarks for AI code reviewers tend to trade off label quality, coverage and real-world representativeness, which leaves teams without a reproducible way to compare systems. ReviewBench is GitHub's attempt to close that gap by modeling its corpus on over 100 million real pull requests on GitHub and weighting pull-request size toward the reviewable middle and tail rather than tiny single-file changes.
Its ground truth is assembled from human reviewers, issues inferred from author follow-up commits, deterministic analysis tools and multiple frontier LLMs, then deduplicated and judged under one published rubric. Reviewer agents are scored on two metric families: grounded precision and recall against existing gold-set labels, and augmented precision and recall that can credit valid findings no gold-set producer surfaced. Results can be sliced by severity and category, or re-ranked by adjusting β in the Fβ score for teams that prefer broader coverage or less noise.
GitHub says it has used the benchmark to evaluate Copilot code review across successive iterations, and that offline changes have consistently pointed in the same direction as later production A/B tests. The stakes therefore hinge on whether that offline-to-online alignment holds for agents built outside GitHub, which the public leaderboard and self-serve runner are designed to test.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Ghost AI raised $11 million, led by Andreessen Horowitz with Abstract, Audacious Ventures, Nova and SV Angel…
OpenAI will roll out invisible text watermarks to all ChatGPT and Codex plan users in the EU within weeks, and…

Reflection announced Beam, a 501B-parameter open-weight model with 23B active parameters, claiming parity with…

A Stanford, Carnegie Mellon, UC Berkeley, and Microsoft Research team ran 6,800+ math, coding, and science tas…

Reflection AI launched Beam, a 501 billion-parameter open-source LLM
A step-by-step guide fine-tunes Muse Glimmer, Meta's 30B vision model, locally for equation-to-LaTeX conversio…
