Agent Arena: DeepSeek V4.1 Flash hits Pareto frontier at $0.06/task with +4.87% net improvement
arena · x · 2026-09-15
Agent Arena published a dynamic leaderboard ranking models on real-world agentic tasks (tool reliability, task completion, steerability), plotted as net improvement vs. cost per task across 1.69M sessions and 43 models.
- Claude Fable 5.1 (Max) tops the chart at +13.90% net improvement but costs $4.54/task; GPT 6 Astra (Max) follows at +11.90%, $4.09/task.
- DeepSeek V4.1-Flash (Max) (MIT-licensed) lands on the Pareto frontier with +4.87% improvement at a median cost of just $0.06/task — the extreme value pick.
- Open-source Tencent Hy4 preview scores +4.96% at $0.22/task; Kimi K3 (Max) posts +6.39% at $0.77/task.
More from coding & agent
- OpenAI's Codex app adds official Arch Linux support via pacman; ChatGPT desktop hits Linux — OpenAIDevs · 2026-09-15
- LangChain's Managed Deep Agents now run from Slack mentions, DMs, and thread replies — Hacubu · 2026-09-15
- Phil Schmid: output schema, statelessness and code mode will drive an MCP comeback — _philschmid · 2026-09-15
- Open-source Hypit lets coding agents clone viral video workflows, ship 100 variants per command — rohanpaul_ai · 2026-09-15
- How the GrokBot design team uses AI agents: Figma Bro bot and voice memos to production code — soleio · 2026-09-15
- LangChain Open-Sources Its Internal Paid Media Agent for Ad Campaigns — Hacubu · 2026-09-15