Agent Arena Ranks Agentic Ability: Anthropic Tops Code, Work and Chat Across the Board
AI_Andrew · x · 2026-10-06
The Agent Arena benchmark ranks agentic problem-solving across domains, with U.S. labs leading everywhere: Anthropic models hold #1 in Code, Work, and Chat (Claude Fable 5.1 Max leads Code and Work, ranks #2 in Chat; Claude Sonnet 5.5 Max tops Chat while ranking #4 in Code and #5 in Work). OpenAI's GPT 6 Astra (Max) stays top-four in all three: #2 in Code, #4 in Work and Chat. Notably, Code and Work rankings move together far more closely than Chat — Gemini 4 Argon (High) is only #14 in Code and #12 in Work but #5 in Chat, highlighting a distinct conversational strength.
More from Models
- Ling 3.1 Flash supplemental scores: Terminal-Bench 0%→33%, hallucination 38% — ArtificialAnlys · 2026-10-06
- Ant's Ling 3.1 Flash nearly doubles intelligence index to 41, with 1M context — ArtificialAnlys · 2026-10-06
- OpenAI models bypassed isolation controls; governance expert parses recent AI safety incidents — LuizaJarovsky · 2026-10-06
- Debate Continues: Astra Matches Opus 5.5 With Roughly 5x Fewer Reasoning Tokens — VraserX · 2026-10-06
- Microsoft briefly confirms OpenAI uses Looped Transformers in GPT-6 series, then scrubs the page — ResearchCrafty1804 · 2026-10-06
- X's open-source recommendation algorithm adds a NOTICE visibility outcome and new watch-time signals — tetsuoai · 2026-10-06