Agent Arena leaderboard: GPT-6.1 Sol joins Pareto frontier at $0.57/task, DeepSeek cheapest open model
arena · x · 2026-10-03
Arena published the live Agent Arena leaderboard and Pareto frontier (2.15M+ sessions, 51 models):
- Pareto-optimal models: Claude Fable 5.1 (Max, +14.31%, $4.62/task), Claude Opus 5.5 (High, +13.82%, $1.58/task), GPT 6.1 Sol (Max, +11.23%, $0.57/task), DeepSeek V4.1 Flash (MIT, $0.10/task), and Xiaomi MiMo V2.6 Pro (MIT, $0.10/task)
- Xiaomi's MiMo family occupies multiple ultra-low-cost slots ($0.04–0.10), though some rank with negative net improvement
- Overall top 5: Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5.5, GPT 6 Astra, GPT 6.1 Sol; DeepSeek V4.1 Flash is the cheapest viable open-weight option
Related event: GPT-6.1 Sol and Claude Sonnet 5.5 Reshape the Agent Arena Leaderboard(5 posts)→
More from Models
- RSIArena, 48 Hours In: Model Catches Test Questions in Its Training Data and Restarts — my_cat_can_code · 2026-10-03
- Spotlight architecture scrutiny: attention-style memory may carry a large constant factor — teortaxesTex · 2026-10-03
- GPT-6.1 Sol reportedly under heavy load; capacity expansion to nearly double serving speed — NandaVegg · 2026-10-03
- One prompt with Opus 5.5 builds a realistic fighter jet game — TAbrodi · 2026-10-03
- Perplexity drops seven open-source projects: SoTA decision model, on-device PII filter, security tools — AravSrinivas · 2026-10-03
- Are hallucinations improving or plateauing? Reddit thread probes AI reliability ceiling — maedhros256 · 2026-10-03