Flagship models cost 2-4x more for marginal gains: GPT-6 Astra at $3.94/task vs $1.03

arena · x · 2026-09-17

LMArena compared the top model from each family by net improvement score and median cost per task.

Key takeaway: GPT-6 Astra and Claude Fable 5.1 stand out — but both cost far more than their labs' other models despite relatively close performance, raising questions about flagship-tier value.

Related event: Agent Arena: GPT-6 Astra Costs Double for Modest Gains(2 posts)→

Original post →

More from Models

Models channel →