UiPath: Haiku's agent success rate halves with every doubling of task length vs Sonnet
EvalRaccoonDev · reddit · 2026-09-22
UiPath's engineering team found Claude Haiku 4.5 only 6 points behind Sonnet on SWE-bench Verified (73.3% vs 79.6%) but massively behind (42.0% vs 85.6%) on their internal agent tasks with hidden checks. Bucketed by command count: tasks of 1-2 commands tie at 89.5%, but at 41+ commands Haiku falls to 21.2% vs Sonnet's 78.8% — every doubling of task length roughly halves Haiku's odds of finishing. The gap is invisible to SWE-bench because its agent can run tests and fix wrong steps, while hidden checks mean one wrong flag at step 3 stays wrong at step 40.
More from coding & agent
- Cloudflare launches Worker Previews: production-like envs for every agent change — dinasaur_404 · 2026-09-22
- Claude as a One-Person Company: A Full Org Chart of AI Skills — Shruti_0810 · 2026-09-22
- Xiaomi's Luo Fuli Recaps MiMo-V2.6 RL Run: 25K Agent Trajectories Per Step — bookwormengr · 2026-09-22
- Cloudflare Worker Previews: one-command deploys, per-PR URLs, isolated state — dinasaur_404 · 2026-09-22
- Worker Previews auto-provisions isolated Durable Objects and containers per preview — dinasaur_404 · 2026-09-22
- Worker Previews config model: branch-level var/secret overrides without touching prod — dinasaur_404 · 2026-09-22