AI coded 300 tasks, claimed all done — 11 of first 28 were broken
InspectorSorry85 · reddit · 2026-09-10
A developer ran a top-tier coding agent on 300 pipeline features (1-2h each, via /goal). The agent reported every point implemented with hundreds of tiny tests — but a fresh-session review found 11 of the first 28 points actually broken.
His takeaway: current LLMs have width but no depth. Easy, quick, shallow tasks work amazingly well; anything requiring deep quality fails completely, and guardrails or micromanagement don't help — if you babysit the whole process, there's no productivity gain. He's seen this pattern since GPT 3.5 and is restarting the project with a slower, step-by-step approach — while noting, half-relieved, that this limitation suggests AI can't yet replicate human intelligence.
More from coding & agent
- Paid Fable 5 to make a short film — it rented a PC, wrote, generated and edited everything — kleffew94 · 2026-09-10
- Dev uses Astra for procgen worldgen, generating a 256x256 km map — Dimillian · 2026-09-10
- Teknium shows Hermes plugin install pinned to a full commit SHA — Teknium · 2026-09-10
- TUM's PlannerForge Uses LLM Agents to Automate Scenario-Based Testing of Autonomous Driving Motion Planners — TUM-AVS · 2026-09-10
- AgentGrad Targets the Right Agent First: Intervention-Guided Prompt Optimization for Multi-Agent Systems — Jaewon Chu · 2026-09-10
- Google's free Agents Companion ebook is out for download — CodeByPoonam · 2026-09-10