Haiku fails 15 of 17 agent runs on output format, deemed no upgrade over Luna

zeeg · x · 2026-10-11

Sentry CEO David Cramer (zeeg) quoted agent benchmark results: of 17 runs, Haiku returned results that didn't fit the required output shape (extra or misplaced fields) in 15, and ran out of its 18-turn budget in 2. He had hoped Haiku might be competitive with Luna as an upgrade, but concluded it likely isn't — highlighting weak structured-output reliability in schema-constrained agent workflows.

Original post →

More from coding & agent

coding & agent channel →