Agentic Benchmark Checklist paper shows flawed agent benchmarks skew results by up to 100%
ddkang · x · 2026-09-19
A 26-author paper from Berkeley/Stanford-adjacent researchers (Percy Liang, Matei Zaharia, Ion Stoica, Daniel Kang among them) shows many agentic benchmarks suffer from flawed task setup and reward design:
- SWE-bench Verified uses insufficient test cases; TAU-bench counts empty responses as successes — such issues can under- or overestimate agent performance by up to 100% in relative terms
- The authors introduce the Agentic Benchmark Checklist (ABC), guidelines synthesized from benchmark-building experience, best-practice surveys, and previously reported issues
- Applying ABC to CVE-Bench, a benchmark with a particularly complex evaluation design, reduces performance overestimation by 33%
The 39-page paper offers an actionable standard for rigorous agent evaluation.
More from coding & agent
- Claude Code now supports AGENTS.md as fallback when CLAUDE.md is absent — jasonkneen · 2026-09-19
- Pydantic Creator Benchmarks Jev vs Sonnet: 2.6x Cheaper, 3.7x Faster in 17 Lines — samuelcolvin · 2026-09-19
- GPT-6 Astra clears Geometry Dash demon level Jumper with all 3 coins — imjustnewatai · 2026-09-19
- Greg Kamradt: Humans with AI still beat AI with AI on productivity — GregKamradt · 2026-09-19
- Veteran ML engineer: Jev may push agent tool-calling back to discriminative models — multiply_matrix · 2026-09-19
- Box CEO demos Jev for instant enterprise document triage at near-zero cost — multiply_matrix · 2026-09-19