Opus 5.5 xhigh aces the Boeing 747 benchmark in hands-on test
victormustar · x · 2026-09-23
victormustar shared results of Claude Opus 5.5 (xhigh) on the Boeing 747 benchmark, calling the performance "excellent as expected" and publishing the prompt and run details. He previously ran Fable 5.1 on the same benchmark and found it clearly better than 5.0. A follow-up notes Opus 5.5's standout trait: building the right tools to continuously improve its own work.
Related event: Opus 5.5 aces Boeing 747 benchmark, praised for building its own tools(2 posts)→
More from Models
- GPT-6 verdicts split: Sol disappoints, Luna surprises, and Opus 5.5 wins users back — xeophon · 2026-09-23
- Meme Roasts Google for Endless Flash Models Instead of Competing at the Top — RichardMortis12 · 2026-09-23
- Decision-only model Jev beats frontier Gemini on 1,759 real decisions: 5x faster, 25x cheaper — Humble_News_6994 · 2026-09-23
- 40% Opus price cut and AGENTS.md support: Anthropic's four plays to win back developers — iannuttall · 2026-09-23
- Anthropic and OpenAI price cuts show open-source models are winning — bindureddy · 2026-09-23
- Apple AI is still unusable: asks for today's date, confidently answers Oct 10, 2024 — bernard_hossmoto · 2026-09-23