GitHub Open-Sources ReviewBench, a Code Review AI Benchmark Built on 100M Real PRs
GitHub Blog AI/ML · rss · 2026-10-05
GitHub released ReviewBench, an open offline benchmark for agentic code review, modeled on the language, repo-size, and PR-size distribution of over 100 million real pull requests. It uses a multi-source golden set, a consistent rubric validated by senior engineers, and supports breakdowns by severity, category, and precision-recall tradeoffs (F1/Fβ). GitHub says ReviewBench-driven offline evals of Copilot code review now better predict production experiment outcomes, and teams can onboard their own reviewers and submit results today.
More from coding & agent
- OpenCode hits 17M MAUs just 16 months after launch — ycombinator · 2026-10-06
- Claude Code 2.1.291 fixes cloud session and message-loss regressions — ClaudeCodeLog · 2026-10-06
- AI agent architecture explained: the 7 modules from perception to observability — ZabihullahAtal · 2026-10-06
- Memoria 1.0: a local, model-agnostic LLM memory layer hitting 89.8% Recall@1 on LongMemEval — kitkatz69 · 2026-10-06
- Dev resurfaces his 4-year-old UE5 AutoLOD tool for optimizing GenAI 3D game assets — rms80 · 2026-10-06
- Marketers ask: how do we pitch offers to users' AI agents instead of users? — iamthebulk · 2026-10-06