Same Spec Through Cursor, Codex and Claude Code: Every Failure Was in the Wiring, Not the Code

SSShken · reddit · 2026-10-02

A developer ran the exact same feature spec through Cursor, Codex, and Claude Code on one real codebase (cheapest paid tier, defaults, scored against a pre-written 16-point checklist).

Core finding: code quality was fine everywhere; failures were all in the wiring. Two of three tools added the requested doc but never wired it into the list that tells agents what to read — that list is hardcoded in four places, one saying "don't reference anything outside this list," which a fresh agent can't know. Codex styled only one of three screens; the promised later styling pass never came, so two screens shipped unstyled.

Environment gotchas: Cursor silently loaded the author's local Claude Code plugins by default; a fresh account didn't help since plugins live in the home folder; the only clean run came from pointing CLAUDECONFIGDIR at an empty dir.

The real differentiator was scope ownership: with the local DB down, Cursor stopped at "couldn't verify"; Codex started Docker itself, ran migrations, wrote Playwright tests, and found a serialization bug in its own code; Claude Code did all that plus a contrast test on presets it had just invented, catching two WCAG AA failures.

Numbers: Claude Code 38 min / 6% of weekly limit; Cursor 52 min / 3% of monthly quota; Codex hit its 5h cap mid-feature, waited 3.5h, then finished. Caveat: one run each. Spec, screenshots, and three live demos in comments.

Related event: Hands-on Comparison: Cursor, Codex and Claude Code Fail at Wiring, Not Code(2 posts)→

Original post →

More from coding & agent

coding & agent channel →