Agent Arena Leaderboard: Claude Opus 5 Leads, Kimi K3 Enters Top 5
arena · x · 2026-08-05
Agent Arena released its latest leaderboard, evaluating 45 models across 1.6 million sessions on real-world agentic tasks. Claude Opus 5 (High) tops with a 12.10% net improvement, followed by Claude Fable 5 (High) and Claude Opus 5 (Max). Kimi K3 (Max) ranks fourth at 10.18%, and GPT-5.6 Sol is fifth at 10.09%. The board includes metrics like tool hallucination rate and steerability.
More from Models
- User Hits Grok Content Restrictions While Trying to Generate Meme Image — arieljalali · 2026-08-05
- DiffusionGemma Report: Parallel 256-Token Generation Breaks AR Bottleneck — SungjinAhn_ · 2026-08-05
- Gemini Flash Hits Its Limit: Cannot Solve IMO-2025 Problem 6 — teortaxesTex · 2026-08-05
- OpenAI's New Model Scores 35% on SimpleQA, Below Pro Preview's 57% — teortaxesTex · 2026-08-05
- AI Agents Show Significant Decline in Instruction Following — RexDouglass · 2026-08-05
- OpenAI's Codex Shows Monopolistic Trend, Squeezing Native Agent Frameworks — vista8 · 2026-08-05