Flagship models cost 2-4x more for marginal gains: GPT-6 Astra at $3.94/task vs $1.03
arena · x · 2026-09-17
LMArena compared the top model from each family by net improvement score and median cost per task.
- OpenAI: GPT-6 Astra (Max) +11.7% at $3.94/task; GPT-5.6 Sol (xHigh) +7.0% at just $1.03/task
- Anthropic: Claude Fable 5.1 (Max) +13.7% at $4.40/task; Claude Opus 5 (High) +10.2% at $2.07/task
Key takeaway: GPT-6 Astra and Claude Fable 5.1 stand out — but both cost far more than their labs' other models despite relatively close performance, raising questions about flagship-tier value.
Related event: Agent Arena: GPT-6 Astra Costs Double for Modest Gains(2 posts)→
More from Models
- YC built AI versions of its partners on GLM-5.2, cutting latency 31% vs OpenAI — ycombinator · 2026-09-17
- Astra isn't GPT-6 itself — it's just one tier alongside sol and luna — flowersslop · 2026-09-17
- Open weights are not open source: why AI's favorite label is under dispute — StanfordHAI · 2026-09-17
- New Class of AI 'Judgment Models' Like Jev Could Reshape Business Automation — The AI Daily Brief · 2026-09-17
- Parallel decoding vs. structured outputs: devs speculate on a closed-source release — ricklamers · 2026-09-17
- fx's safety classifier benchmarked: ~5-18x faster and more accurate than GPT-5.6-Luna — cramforce · 2026-09-17