Tested Cursor, Claude Code, Codex and Antigravity on the Same App Build, Bugs Included
workflowverdict · reddit · 2026-09-17
The author ran an identical-prompt, identical-rules comparison of Cursor, Claude Code, Codex and Antigravity on the same app build. Episode 1: results were much closer than expected, with no clear winner. Episode 2: they upped the difficulty by giving all four agents a broken production app containing 12 bugs, then validated each agent's fixes against hidden tests the agents couldn't see, filtering out superficial patches.
More from coding & agent
- Dev generates a full software promo video in pure code with GPT-6 Astra, no generative models — op7418 · 2026-09-17
- Frida, an iOS Personal Life Agent, Opens Second TestFlight Batch — Scobleizer · 2026-09-17
- Doubao Seed-2.1-pro-0915 hands-on: tool-calling score jumps 38.4% to 64.8%, coding cost down 40% — vista8 · 2026-09-17
- Dev claims harness makes agent computer use consume identical tokens to regular tool use — TejasKumar_ · 2026-09-17
- Spreadsheet agents should deliver editable drafts, not just approval summaries — Mountain_Athlete1350 · 2026-09-17
- Score Studio Launches as All-in-One Vision AI Platform for Annotation, Training and Deployment — markjeffrey · 2026-09-17