Claude Haiku 5.5 ships with huge jumps: OSWorld 15.7%→72.4%, beats GPT-6 Luna across the board

mark_k · x · 2026-10-08

Anthropic's smallest model Claude Haiku 5.5 just dropped with a massive benchmark leap, beating GPT-6 Luna on every chart where both are reported: OSWorld 2.1 computer use 15.7%→72.4% (vs Luna's 48.9%), HLE without tools 10.2%→45.9% (57.4% with tools), Terminal-Bench 4.0 0%→39.2% (vs Luna's 16.4%), FrontierCode 46.4%, Chartography 6.4%→46.4%, and knowledge-work scores of 1620 GDPval-AA / 1578 AA-Briefcase. Sonnet 5.5 still leads overall, but the generational jump — especially in computer use and reasoning — is enormous.

Related event: Anthropic launches Claude Haiku 5.5 with 75% lower costs(27 posts)→

Original post →

More from Models

Models channel →