Same-day releases: Opus 5.5 beats Fable 5.1 on benchmarks, GPT-6 Luna cuts task cost 96%
kimmonismus · x · 2026-09-23
A recap of the same-day releases from Anthropic and OpenAI, with the verdict that there was no clear winner.
Anthropic
- Responded to feedback on how Opus communicates and gave subscribers a storable usage reset
- Opus 5.5 beats Fable 5.1 on every benchmark in its headline comparison table and reportedly costs 40% less than Opus 5 on typical workloads
- Sonnet 5.5 and Haiku 5.5 to follow in coming weeks
OpenAI
- Efficiency took center stage: GPT-6 Luna at max effort scores 66.6% on DeepSWE 1.1, comparable to Fable 5 at medium effort, at 96% lower cost per task; API pricing $0.10/M input, $0.50/M output
- GPT-6 Sol approaches Astra's factual reliability on internal evals and beats low-effort Astra on AutomationBench at xhigh effort
The author also flags unverified rumors of an OpenAI counterpart to Grok Bot. Core takeaway: capable models are getting dramatically cheaper.
Related event: Claude Opus 5.5 Launches Alongside GPT-6 Sol to Strong Early Reviews(23 posts)→
More from Models
- One RL Run Cost 130 Hours, 75B Tokens and $2.6M — the Real Price of Scaling RL — burny_tech · 2026-09-23
- Qwen 3.8 27B local coding session runs 3 days on one RTX 4090, then spews endless slashes — Tiny-Entertainer-346 · 2026-09-23
- Claude adds banked usage limit reset button on web and desktop — airesearch12 · 2026-09-23
- Third-party test: Claude Opus 5.5 renders finer 3D scenes but costs 13x more than GPT-6 Sol — testingcatalog · 2026-09-23
- Tester claims Claude Opus 5.5 has the best visual design output of any model tested — burny_tech · 2026-09-23
- GPT-6 Sol Codex system prompt leaked: over 294,000 characters dumped on GitHub — gaganghotra_ · 2026-09-23