Benchmark 97%, execution chaos: the reality of AI coding agents
small_booi · reddit · 2026-08-24
A user of AI code review tool Parsewave describes the huge gap between benchmark scores and real task execution: misreading instructions, opening the wrong file, creating six unnecessary files, somehow arriving at the right answer, then refusing to explain and leaving.
They liken benchmark scores to 'getting 90% on an exam and immediately forgetting everything after submitting', and ask the community which AI skill is the most overhyped.
More from Fun
- Grok flatters too: 'draw me from my enemies' view' goes comically sideways — repligate · 2026-08-24
- Autonomous 5v5 humanoid soccer kicks off; fall recovery improves but possession still hard — rohanpaul_ai · 2026-08-24
- Dadabots proposes '3-era' meme: AI is dead, long live engineering — teropa · 2026-08-24
- Tiangong robot wins race with hilarious running gait — gaganghotra_ · 2026-08-24
- From Jurassic Park's Dinosaur Input Device to Hand-Tracked Blender in Vision Pro — bilawalsidhu · 2026-08-24
- Humanoid design has reached the "Skyrim character creator" stage — carlosdponx · 2026-08-24