Opus 5.5 Beats GPT-6 Astra on Most Benchmarks, but Astra Leads Terminal-Bench-Science
BenBlaiszik · x · 2026-09-23
Anthropic positions Opus 5.5 as a major step up from Opus 5, leading on agentic coding, computer use, and knowledge work. Third-party comparison shows Opus 5.5 outperforms Fable 5.1 and GPT-6 Astra on most benchmarks, though Astra retains a clear lead on Terminal-Bench-Science — the poster asks for hypotheses on why.
Related event: Claude Opus 5.5 Sweeps Benchmarks, Tops AA Intelligence Index(78 posts)→
More from Models
- OpenAI's new internal model reportedly solved 100+ open math problems in under 24 days — haider1 · 2026-09-23
- Reddit user on Opus 5.5: coding is basically solved, it's absurd — rocket_zen · 2026-09-23
- Rumor mill: some claim Opus 3.5 became Opus 4, others suspect it's really an 'Opus 3.6' — repligate · 2026-09-23
- Developer calls new LLM comparison charts 'practically worthless' as everyone tests differently — jdluk87 · 2026-09-23
- Matt Shumer: using Opus without hitting rate limits feels 'freeing' — mattshumer_ · 2026-09-23
- SVG showdown: Opus 5.5 (High) vs GPT-6 Sol (High) on the same prompt — OriginalScrubLord · 2026-09-23