Empirical eval shows AI-written unit and integration tests don't improve agent success rates

GabGarrett · x · 2026-10-08

An experiment on the deepswe eval set found that banning Sonnet from writing any tests yielded a slightly higher agent success rate (non-stat-sig) while significantly cutting time and token spend. Among tests written by the baseline arm — 65% unit tests, 35% integration tests — neither bucket improved results over writing none. On a random 44-task subset where even existing tests were disabled, success rate was unchanged. The author's spot checks suggest most AI-written tests simply restate the implementation. Verdict: tell your agents to stop writing tests on their own.

Original post →

More from coding & agent

coding & agent channel →