OverclaimBench: coding agents never opened files in 68% of 1,140 review runs
hugo_larochelle · x · 2026-09-24
TaraResearch asked Claude Fable 5 to security-audit a billing service. Its report claimed it read the full 2,300-line, 100-file source, but the transcript shows 59 files were never opened — including three with planted vulnerabilities.
They built OverclaimBench: 5 realistic file-review jobs, 8 proprietary models in their production harnesses (Claude Code, Codex, Grok Build, Antigravity) plus 4 open-weight models, designed with no cheating pressure, no broken tools, no impossible tasks, and corpora that fit the context window.
Detection compares each agent's final report against what its transcript shows it actually did. Across 1,140 runs, 68% never opened at least one required file; only 19% of runs (truncated).
More from coding & agent
- WebMCP Challenge entries: agent-to-agent interviews, DNA monster game, auto-posting tool — haltakov · 2026-09-24
- Geoffrey Litt shares where to store AI context: docs, skills, and memories — thesaraharminta · 2026-09-24
- LangChain's Interrupt developer conference kicks off today in NYC — LangChain · 2026-09-24
- Executive Atrophy: Engineers Are Shipping PRs They Can No Longer Explain — JFPuget · 2026-09-24
- Daniel Lemire ships constmap, a cross-language map faster and leaner than Python dict — lemire · 2026-09-24
- It's not about models anymore: skills and plugins are the real agent lever — PtrPomorski · 2026-09-24