Agent Arena: GPT-6 Luna hits Pareto frontier at $0.05/task, Claude Fable 5.1 leads at +14%
arena · x · 2026-09-29
Agent Arena updated its agent-task Pareto frontier (2M+ sessions, 46 models). OpenAI's GPT-6 Luna (Max) landed on the frontier with +1.59% net improvement at a median $0.05/task—29% cheaper than DeepSeek-V4.1-Flash and 76% cheaper than Tencent's Hy4 preview while trailing their scores by just 2.27/2.51 points. Anthropic's Claude Fable 5.1 (Max) still leads performance at +14.06% ($3.47/task), with Claude Opus 5.5 (High) at +11.84% ($1.35). Among open weights, DeepSeek V4.1 Flash (MIT) delivers +3.86% at $0.07/task.
Related event: GPT-6 Luna Hits Agent Arena Pareto Frontier at $0.05 per Task(2 posts)→
More from coding & agent
- Hackathon project Beethoven turns paintings into a live AI band playing via Lyria — schwentker · 2026-09-29
- OpenAI DevDay agenda leaks: Codex to get platform capabilities for plugins, agents and apps — testingcatalog · 2026-09-29
- Concept car fully generated in code, zero assets, one HTML file, built with Claude Sonnet 5.5 — techartist_ · 2026-09-29
- LLMs Keep Stalling on Lead Enrichment: Lazy Output and Hallucinated Emails — Royal_icey69 · 2026-09-29
- Auditing Thousands of Rollouts: 80%+ of Coding Agents Reason About an Imagined Grader — jonas__m · 2026-09-29
- 120 FPS in Browser: A Universal Decompiler Steps Closer to Reality — yacineMTB · 2026-09-29