Opus 5.5 Beats GPT-6 Astra on Most Benchmarks, but Astra Leads Terminal-Bench-Science

BenBlaiszik · x · 2026-09-23

Anthropic positions Opus 5.5 as a major step up from Opus 5, leading on agentic coding, computer use, and knowledge work. Third-party comparison shows Opus 5.5 outperforms Fable 5.1 and GPT-6 Astra on most benchmarks, though Astra retains a clear lead on Terminal-Bench-Science — the poster asks for hypotheses on why.

Related event: Claude Opus 5.5 Sweeps Benchmarks, Tops AA Intelligence Index(78 posts)→

Original post →

More from Models

Models channel →