Empirical data: banning AI-written unit tests slightly boosts agent success

GabGarrett · x · 2026-10-09

On the deepswe eval set, forbidding Sonnet from writing any tests slightly raised agent success rates (not stat-sig) while significantly cutting time and tokens (stat-sig). Of baseline-written tests, 65% were unit and 35% integration — neither improved results. The author also ran a 44-task subset with existing tests disabled. Shaw confirms he switched to e2e-only testing months ago for faster, higher-quality iteration.

Related event: Banning AI-Written Unit Tests May Actually Boost Agent Efficiency(2 posts)→

Original post →

More from coding & agent

coding & agent channel →