GPT-6 Astra vs Opus 5.5: Opus Wins Knowledge Work, OpenAI Keeps Hard STEM

johnseach · x · 2026-09-23

A detailed head-to-head. On the Artificial Analysis Intelligence Index, Opus 5.5 leads at 58, with Astra and Fable 5.1 at 53 and old Opus 5 at 51.

Pricing isn't close: Opus 5.5 is $4/$20 per million tokens vs Astra's $10/$50; cache reads are $0.20 vs $1.00, making Opus 60% cheaper. But Astra is extremely token-thrifty (27k vs 119k tokens per task), so real bills are closer than sticker prices suggest.

Coding is murky: Anthropic's own table shows Opus 5.5 at 66.4% vs 57.9% on Terminal-Bench 4.0, but Artificial Analysis, running both in the same harness, has them roughly tied at 59–60%. Vendor charts flatter the home team; trust the independent run.

Knowledge work is Opus 5.5's clean win: GDPval-AA at 1846 Elo vs 1542, Humanity's Last Exam with tools at 67.7% vs 57.2%. Long reports, repo migrations, and full-brief agents favor Opus.

Astra owns hard STEM and cyber: FrontierMath Tier 4 at 97.6%, ExploitBench at 100% (now past OpenAI's Critical cyber threshold, so gated), ScreenSpot-Pro at 92.7%.

Bottom line: pick Opus 5.5 as the daily driver; keep Astra for frontier math, terminal science, screen-heavy use, and gated cyber. OSWorld scores aren't comparable due to differing scoring rules. Specs are close (1M/1.05M context, 128K output).

Related event: Claude Opus 5.5 Launches, Tops Intelligence Index While Cutting Prices(50 posts)→

Original post →

More from Models

Models channel →