Opus 5.5 tops AI index at 58, undercuts GPT-6 Astra by 60% but burns 4x more tokens
johnseach · x · 2026-09-23
Anthropic's Claude Opus 5.5 and OpenAI's GPT-6 Astra (1M-class context both) make very different bets.
- Index: Opus 5.5 ranks #1 on Artificial Analysis at 58 vs Astra's 53; leads 6 of 10 tests, biggest on knowledge work and long agent jobs.
- Price: Opus $4/$20 vs Astra $10/$50; cache reads $0.20 vs $1.00 — 60% cheaper tokens, 5x cheaper cache. But Astra is far thriftier (27k vs 119k tokens per task), so real bills are closer.
- Coding: vendor tables conflict (Anthropic claims 66.4% vs 57.9% on Terminal-Bench 4.0); independent same-harness runs show a tie around 59-60%.
- Opus wins knowledge work: GDPval-AA 1846 vs 1542 Elo; HLE with tools 67.7% vs 57.2%.
- Astra keeps STEM/cyber: FrontierMath Tier 4 at 97.6%, ExploitBench 100% (gated at OpenAI's Critical cyber threshold), ScreenSpot-Pro 92.7%. OSWorld scores aren't comparable across scoring rules.
Verdict: Opus 5.5 as the daily driver; keep Astra for frontier math, terminal science, screen-heavy computer use, and gated cyber.
Related event: Claude Opus 5.5 Launches, Tops Intelligence Index While Cutting Prices(50 posts)→
More from Models
- Third-party test: Claude Opus 5.5 renders finer 3D scenes but costs 13x more than GPT-6 Sol — testingcatalog · 2026-09-23
- GPT-6 Sol priced at half of Opus 5.5 as Sol and Luna go 'dirt cheap' — ZeroStateReflex · 2026-09-23
- Tester claims Claude Opus 5.5 has the best visual design output of any model tested — burny_tech · 2026-09-23
- Meta's Alexandr Wang reveals muse has been in the works since at least Sept 2025 — adrianscottcom · 2026-09-23
- GPT-6 Sol Codex system prompt leaked: over 294,000 characters dumped on GitHub — gaganghotra_ · 2026-09-23
- Claude 5.5 (live) keeps generating user turns, reports user — BlackHC · 2026-09-23