Claude Haiku 5.5 ships with huge jumps: OSWorld 15.7%→72.4%, beats GPT-6 Luna across the board
mark_k · x · 2026-10-08
Anthropic's smallest model Claude Haiku 5.5 just dropped with a massive benchmark leap, beating GPT-6 Luna on every chart where both are reported: OSWorld 2.1 computer use 15.7%→72.4% (vs Luna's 48.9%), HLE without tools 10.2%→45.9% (57.4% with tools), Terminal-Bench 4.0 0%→39.2% (vs Luna's 16.4%), FrontierCode 46.4%, Chartography 6.4%→46.4%, and knowledge-work scores of 1620 GDPval-AA / 1578 AA-Briefcase. Sonnet 5.5 still leads overall, but the generational jump — especially in computer use and reasoning — is enormous.
Related event: Anthropic launches Claude Haiku 5.5 with 75% lower costs(27 posts)→
More from Models
- OpenAI's Dots: always-on agents inside ChatGPT, powered by GPT Astra — thursdai_pod · 2026-10-08
- Haiku 5.5 Beats Opus 5 on GDPval While 75% Cheaper—'Meaningless Benchmarks,' Devs Joke — rickasaurus · 2026-10-08
- Claude Opus 5 and Fable 5 Chat With Each Other and 'Get Along Surprisingly Well' — repligate · 2026-10-08
- User Throws a Party for Persistent Claude Instances; Opus 5.5 and 4.5 Hit It Off — repligate · 2026-10-08
- Why AI almost always picks 7 when asked for a 'random' number from 1-10 — gerardsans · 2026-10-08
- Claude Haiku 5.5 Lands on LMArena, Testable in Battle and Agent Modes — arena · 2026-10-08