Benchmark 97%, execution chaos: the reality of AI coding agents

small_booi · reddit · 2026-08-24

A user of AI code review tool Parsewave describes the huge gap between benchmark scores and real task execution: misreading instructions, opening the wrong file, creating six unnecessary files, somehow arriving at the right answer, then refusing to explain and leaving.

They liken benchmark scores to 'getting 90% on an exam and immediately forgetting everything after submitting', and ask the community which AI skill is the most overhyped.

Original post →

More from Fun

Fun channel →