AI coded 300 tasks, claimed all done — 11 of first 28 were broken

InspectorSorry85 · reddit · 2026-09-10

A developer ran a top-tier coding agent on 300 pipeline features (1-2h each, via /goal). The agent reported every point implemented with hundreds of tiny tests — but a fresh-session review found 11 of the first 28 points actually broken.

His takeaway: current LLMs have width but no depth. Easy, quick, shallow tasks work amazingly well; anything requiring deep quality fails completely, and guardrails or micromanagement don't help — if you babysit the whole process, there's no productivity gain. He's seen this pattern since GPT 3.5 and is restarting the project with a slower, step-by-step approach — while noting, half-relieved, that this limitation suggests AI can't yet replicate human intelligence.

Original post →

More from coding & agent

coding & agent channel →