Empirical data: banning AI-written unit tests slightly boosts agent success
GabGarrett · x · 2026-10-09
On the deepswe eval set, forbidding Sonnet from writing any tests slightly raised agent success rates (not stat-sig) while significantly cutting time and tokens (stat-sig). Of baseline-written tests, 65% were unit and 35% integration — neither improved results. The author also ran a 44-task subset with existing tests disabled. Shaw confirms he switched to e2e-only testing months ago for faster, higher-quality iteration.
Related event: Banning AI-Written Unit Tests May Actually Boost Agent Efficiency(2 posts)→
More from coding & agent
- anti-slop: an open-source rulebook with 5.1k stars to strip generic AI slop from coding agents — tom_doerr · 2026-10-09
- IBM's RIT-RAG induces document sub-trees from retrieved chunks, lifting RAG accuracy by up to 11.4 points — _reachsumit · 2026-10-09
- Reddit thread rounds up open-source coding agents like MonkeyCode and OpenHands for local use — Good_Full_Tmes · 2026-10-09
- OpenAI's 10,000-agent, 130B-token run pushed slime v0.4.0 to rethink RL infrastructure scale — teortaxesTex · 2026-10-09
- Voice agent stack breakdown: Twilio, Deepgram, Cresta and ElevenLabs with a 300ms latency rule — schwentker · 2026-10-09
- Refund tools as verbs, not pens: how 'Veronica' stops LLM agents from inventing payees — schwentker · 2026-10-09