Same Spec Through Cursor, Codex and Claude Code: Every Failure Was in the Wiring, Not the Code
SSShken · reddit · 2026-10-02
A developer ran the exact same feature spec through Cursor, Codex, and Claude Code on one real codebase (cheapest paid tier, defaults, scored against a pre-written 16-point checklist).
Core finding: code quality was fine everywhere; failures were all in the wiring. Two of three tools added the requested doc but never wired it into the list that tells agents what to read — that list is hardcoded in four places, one saying "don't reference anything outside this list," which a fresh agent can't know. Codex styled only one of three screens; the promised later styling pass never came, so two screens shipped unstyled.
Environment gotchas: Cursor silently loaded the author's local Claude Code plugins by default; a fresh account didn't help since plugins live in the home folder; the only clean run came from pointing CLAUDECONFIGDIR at an empty dir.
The real differentiator was scope ownership: with the local DB down, Cursor stopped at "couldn't verify"; Codex started Docker itself, ran migrations, wrote Playwright tests, and found a serialization bug in its own code; Claude Code did all that plus a contrast test on presets it had just invented, catching two WCAG AA failures.
Numbers: Claude Code 38 min / 6% of weekly limit; Cursor 52 min / 3% of monthly quota; Codex hit its 5h cap mid-feature, waited 3.5h, then finished. Caveat: one run each. Spec, screenshots, and three live demos in comments.
Related event: Hands-on Comparison: Cursor, Codex and Claude Code Fail at Wiring, Not Code(2 posts)→
More from coding & agent
- Dev argues AI is great at assets and code but bad at designing game mechanics that feel good — rms80 · 2026-10-03
- Free one-day curriculum takes you from AI agent basics to MCP and agent security — ifioknkem · 2026-10-03
- Using System One models in Swift: fast, deterministic decisions via Apple Foundation Models — rxwei · 2026-10-03
- Run 50 AI Coding Agents in Parallel With One Global Rule for Background Tasks — Daniel_Farinax · 2026-10-03
- Why Linear's Agent Session beats Slack as a collaboration surface for agentic work — jeff_weinstein · 2026-10-03
- SWE-chat V2 ships 3.5x larger with agent skills and subagent trajectories — Diyi_Yang · 2026-10-03