Tested Cursor, Claude Code, Codex and Antigravity on the Same App Build, Bugs Included

workflowverdict · reddit · 2026-09-17

The author ran an identical-prompt, identical-rules comparison of Cursor, Claude Code, Codex and Antigravity on the same app build. Episode 1: results were much closer than expected, with no clear winner. Episode 2: they upped the difficulty by giving all four agents a broken production app containing 12 bugs, then validated each agent's fixes against hidden tests the agents couldn't see, filtering out superficial patches.

Original post →

More from coding & agent

coding & agent channel →