Agent Arena Pareto frontier: DeepSeek V4.1 Flash delivers +4.1% at just $0.06 per task
arena · x · 2026-09-26
Agent Arena's cost-vs-performance Pareto frontier maps net improvement against median cost per task:
Pareto-optimal tier: Claude Fable 5.1 (Max) +13.80%/$4.11, GPT 6 Astra (Max) +10.85%/$2.82, Claude Opus 5 (High) +9.80%/$2.15, Claude Fable 5 (High) +8.28%/$1.71, GPT 6 Sol (Max) +7.68%/$0.74, Claude Sonnet 5 (High) +4.79%/$0.70.
The budget tier is dominated by Chinese models: Kimi K3 (Max) +4.55%/$0.67, Tencent Hy4 preview +4.37%/$0.17, DeepSeek V4.1 Flash (Max) +4.10%/$0.06, plus Tencent Hy3 and Xiaomi Mimo V2.5 Pro at $0.04 with negative net improvement. Gemini 3.8 Flash and GLM 5.2 also make the top 15.
More from coding & agent
- One agent writes the fix, another reviews it: a two-agent code review workflow in Slack — Al_Grigor · 2026-09-26
- Academic agent Memex upgraded to Opus 5.5: writing quality fixed, experience much better — arjunrajlab · 2026-09-26
- Dev uses open-source Ling-3.0-flash-VL to let AI redesign the foldable iPhone in a single HTML file — alifcoder · 2026-09-26
- Anthropic launches Claude plugin directory portal as MCP usage jumps 110x this year — ClaudeDevs · 2026-09-26
- Open-source Jev agent plays Pokemon Red live, pushing fast-decision AI beyond Tetris — supportingthedogs · 2026-09-26
- Anthropic deep dive: effort tuning in Claude Code pays off most for security and code review — trq212 · 2026-09-26