Agent Arena Pareto frontier page live: 57 models ranked by net improvement vs. cost per task
arena · x · 2026-09-03
The LMArena team shared the Agent Arena Pareto frontier page: a dynamic ranking of 57 models by net improvement × cost per task on real Agent Mode tasks, across 2.17M+ sessions, with signals including tool reliability, task completion, and steerability. This is the follow-up link post to the Gemini 3.8 Flash leaderboard news.
Related event: Agent Arena launches Pareto frontier leaderboard for 57 models(2 posts)→
More from Models
- Fable 5.1 burns 102K reasoning tokens and hits the 128K output ceiling mid-code — rohanpaul_ai · 2026-09-03
- Google ships five models in one week: Gemini 3.8 Flash, Muse Spark 1.3, and more incoming — sethlazar · 2026-09-03
- Every's writing bench adds Gemini 3.8 Flash, Grok 4.6, and Muse Spark 1.3 — danshipper · 2026-09-03
- Meta's Muse Spark 1.3 lands on OpenRouter with 1M context for agentic workflows — armand_ruiz · 2026-09-03
- DeepSeek-V4-Pro ships with 1.6T-param MoE; open-source eval harness steals the show — DeepLearningAI · 2026-09-03
- Rival AI agents: cross-vendor model review catches what self-review misses — rseroter · 2026-09-03