New Arena Leaderboard: Claude Fable 5.1 Tops Agents at 13.71%, Tasks From $0.04
arena · x · 2026-09-25
Arena's redesigned leaderboard (354M+ total sessions) shows: Claude Fable 5.1 (Max) tops overall agents at 13.71%, ahead of GPT 6 Astra (11.54%) and Claude Opus 5 High (10.25%); Claude Opus 5.5 is #1 in WebDev·Max. The Pareto frontier spans Claude Fable 5.1 ($4.17/task) down to Mimo V2.5 Pro and Hy3 at $0.04/task (5%), with Deepseek V4.1 Flash at $0.06. Live sessions feature Gemini 3.8, GPT 5.6, Grok 4.7, GLM 5.3 and anonymous entrants; a new Model Capabilities section offers the Arena team's first impressions, e.g. of GPT-6 Sol.
Related event: Arena Launches Redesigned Leaderboard Overview(2 posts)→
More from Models
- Apple drops new HF model: Qwen3.5-9B finetune that turns long docs into page images to save tokens — yoobinray · 2026-09-25
- Alexandr Wang amplifies Muse skepticism post calling the model 'underwhelming and useless' — maheen_sohail · 2026-09-25
- The AI Meme: Wait Until Your Project Is 78% Done, Then They Nerf the Model — BLUECOW009 · 2026-09-25
- Open Weights Beat Black-Box APIs on Security: Sandboxing Is an Engineering Problem — ypatil125 · 2026-09-25
- Zyphra open-sources ZUNA1.1 EEG foundation model, advancing noninvasive thought-to-text — burny_tech · 2026-09-25
- Contrastive-LM releases CLM-v0.1-8B, a Qwen3-8B-based reranker trending on Hugging Face — Contrastive-LM · 2026-09-25