Four spec-driven coding tools tested on one real feature: 13, 24, 29 tasks and 84 Allow clicks
SSShken · reddit · 2026-09-19
After criticism that his earlier comparison only measured onboarding, the author built a real TypeScript/Express/SQLite/React project with plain Claude Code, gave identical copies to OpenSpec, Spec Kit, BMAD and Kiro, and asked each to add the same one-line feature — with a planted trap: the server already supported it.
- All four shipped correctly and none touched the server or invented endpoints; all read the code before planning
- They diverged elsewhere: task lists of 13, 24, 29, and none; one asked before adding 3 dev dependencies, another added 7 silently; two left test data in the database; one committed straight to main; one required 84 manual Allow clicks and most of a free credit tier
- Everything counted by hand, every diff read manually; the missing control run (plain Claude Code, no workflow) comes next
Related event: Same prompt, four spec-driven coding tools: 4 minutes to 2 hours(2 posts)→
More from coding & agent
- Three coding habits that may never come back in the AI-agent era — BLUECOW009 · 2026-09-20
- mcp-proxmox lets you manage Proxmox VE clusters with natural language — modelcontextprotocol · 2026-09-20
- Tronsave MCP Testnet lets agents buy and sell TRON resources via one interface — modelcontextprotocol · 2026-09-20
- Open questions on JEV: model size, calibration reliability, and a future fine-tuning API — vykthur · 2026-09-20
- One astra general directs 7 jev agents in Warcraft 3: 1,000+ LLM calls for ~20 cents — Vjeux · 2026-09-20
- Composite system1+system2 agent harness runs WC3 micro in RL env, open-sourcing soon — Vjeux · 2026-09-20