Haiku fails 15 of 17 agent runs on output format, deemed no upgrade over Luna
zeeg · x · 2026-10-11
Sentry CEO David Cramer (zeeg) quoted agent benchmark results: of 17 runs, Haiku returned results that didn't fit the required output shape (extra or misplaced fields) in 15, and ran out of its 18-turn budget in 2. He had hoped Haiku might be competitive with Luna as an upgrade, but concluded it likely isn't — highlighting weak structured-output reliability in schema-constrained agent workflows.
More from coding & agent
- Stop Calling Yourself 'Just a Vibe Coder': Learn Enough That Agents Can't Fool You — brandon_galang · 2026-10-11
- Training inside the harness lifts Qwen3-14B from 22.2% to 54.8% on Spider 2.0-SQLite — omarsar0 · 2026-10-11
- Every Ditched Personal Agents: Dan Shipper's OpenClaw 'Died' and No One Restarted It — every · 2026-10-11
- This founder runs his entire product distribution from an Obsidian vault pointed at Claude Code — EXM7777 · 2026-10-11
- RSIGym pattern: keep agents in lightweight CPU containers, offload training/inference/evals to services — SucceededMind · 2026-10-11
- One settings change turns Claude Code into a multi-model team: Opus plans, Sonnet codes, Haiku searches — chessbuzz · 2026-10-11