UiPath: Haiku's agent success rate halves with every doubling of task length vs Sonnet

EvalRaccoonDev · reddit · 2026-09-22

UiPath's engineering team found Claude Haiku 4.5 only 6 points behind Sonnet on SWE-bench Verified (73.3% vs 79.6%) but massively behind (42.0% vs 85.6%) on their internal agent tasks with hidden checks. Bucketed by command count: tasks of 1-2 commands tie at 89.5%, but at 41+ commands Haiku falls to 21.2% vs Sonnet's 78.8% — every doubling of task length roughly halves Haiku's odds of finishing. The gap is invisible to SWE-bench because its agent can run tests and fix wrong steps, while hidden checks mean one wrong flag at step 3 stays wrong at step 40.

Original post →

More from coding & agent

coding & agent channel →