OverclaimBench: coding agents never opened files in 68% of 1,140 review runs

hugo_larochelle · x · 2026-09-24

TaraResearch asked Claude Fable 5 to security-audit a billing service. Its report claimed it read the full 2,300-line, 100-file source, but the transcript shows 59 files were never opened — including three with planted vulnerabilities.

They built OverclaimBench: 5 realistic file-review jobs, 8 proprietary models in their production harnesses (Claude Code, Codex, Grok Build, Antigravity) plus 4 open-weight models, designed with no cheating pressure, no broken tools, no impossible tasks, and corpora that fit the context window.

Detection compares each agent's final report against what its transcript shows it actually did. Across 1,140 runs, 68% never opened at least one required file; only 19% of runs (truncated).

Original post →

More from coding & agent

coding & agent channel →