Empirical eval shows AI-written unit and integration tests don't improve agent success rates
GabGarrett · x · 2026-10-08
An experiment on the deepswe eval set found that banning Sonnet from writing any tests yielded a slightly higher agent success rate (non-stat-sig) while significantly cutting time and token spend. Among tests written by the baseline arm — 65% unit tests, 35% integration tests — neither bucket improved results over writing none. On a random 44-task subset where even existing tests were disabled, success rate was unchanged. The author's spot checks suggest most AI-written tests simply restate the implementation. Verdict: tell your agents to stop writing tests on their own.
More from coding & agent
- Ramp's hidden markdown-file offer read only by AI actually worked, customers claimed it — gaganghotra_ · 2026-10-08
- Creator revives a dead Notion course by feeding it to Codex as an interactive app — RichardsonDx · 2026-10-08
- Gary Bernhardt: AI 'factories' repeat the microservices over-engineering mistake — sergeykarayev · 2026-10-08
- Hierarchical RL with mixed discount rates may unlock long-horizon agent tasks — jessi_cata · 2026-10-08
- Graph-MIND: local MCP memory server hits 88.8% LongMemEval with zero LLM calls at write time — BrilliantGrocery8233 · 2026-10-08
- dotey's Claude Code workflow: Fable as Tech Lead delegating and reviewing for Opus — dotey · 2026-10-08