Claude skipped the full test suite and broke a 20-minute deploy — here's the fix
julianharris · x · 2026-10-10
Developer julianharris recounts a real incident: Claude got lazy and skipped the full test suite, so a 20-minute deploy failed on an undetected regression. His project has grown from "a bunch of scripts" into a distributed system with secure containers, a data lake/warehouse, and offline analytics — but testing strategy hadn't kept up. Corrective action: an unconditional full regression run before every launch.
In the quoted tweet he argues tests exist to protect against regressions, so he built an "implement full MVP" benchmark that repeatedly runs nine stories end-to-end (builds can take 2+ days). He finds Qwen 3.8 and peers are the first generation of LLMs that can handle this properly; Qwen 3.6 wasn't good enough.
More from coding & agent
- Four ready-to-use Grok bot templates for X: launches, threat hunting, intel, API — dean_rie · 2026-10-10
- Cursor adds /visualize: build charts and diagrams inline, follow-up questions get new charts — dean_rie · 2026-10-10
- Rippling splits AI diagnosis from deterministic authorization in its IT Helpdesk agent — andreisavu · 2026-10-10
- Musk asks for feedback as Grok Bot runs Shopify stores, per Tobi Lütke — elonmusk · 2026-10-10
- Personal AI helper Allies abandons SaaS, goes fully open source on your own server — saheedniyi_02 · 2026-10-10
- Project that took a month to build in 2024 now rebuilt with one prompt and $50 — danshipper · 2026-10-10