GitHub Open-Sources ReviewBench, a Code Review AI Benchmark Built on 100M Real PRs

GitHub Blog AI/ML · rss · 2026-10-05

GitHub released ReviewBench, an open offline benchmark for agentic code review, modeled on the language, repo-size, and PR-size distribution of over 100 million real pull requests. It uses a multi-source golden set, a consistent rubric validated by senior engineers, and supports breakdowns by severity, category, and precision-recall tradeoffs (F1/Fβ). GitHub says ReviewBench-driven offline evals of Copilot code review now better predict production experiment outcomes, and teams can onboard their own reviewers and submit results today.

Original post →

More from coding & agent

coding & agent channel →