Kimi K3 Max Tops Arena Leaderboard in Frontend Code and Agent Tasks
arena · x · 2026-07-28
In the latest Arena leaderboard update, Moonshot's Kimi K3 (Max) takes the #1 spot overall and among open-weight models.
- Frontend Code: Ranks #1 among open-weight models across all 7 domains. Overall, it secures #1 in 5 out of 7 domains, landing #2 only in Gaming and Content Creation Tools.
- Agent Arena: Takes the #1 spot with a +9.75% net-improvement, surpassing GLM-5.2 (Max) at +7.12%.
- Evaluation Methodology: Agent Arena measures models on real-world, long-horizon agentic tasks by giving them access to web search, filesystem, and terminal tools. It uses causal tracing methodology to measure the model's net improvement to outcomes.
Related event: Kimi K3 Tops Agent Arena Leaderboard Among Open-Weight Models(4 posts)→
More from Models
- Anthropic rumors point to a larger internal teacher model and a near-K3 public stack — teortaxesTex · 2026-07-28
- Kimi K3’s license and pricing make the model far less usable, users say — Eyelbee · 2026-07-28
- Kimi K3 report details kernel tuning, a Triton-like compiler, and a chip prototype — stochasticchasm · 2026-07-28
- Kimi report reveals a wide internal benchmark suite for coding and agent skills — stochasticchasm · 2026-07-28
- Claude is still being called the most steerable model set, despite its weirdness — sloppenheimer · 2026-07-28
- Kimi K3 lands as a 2.8T open model with a 1M-token context window — Ambitious_Ad4397 · 2026-07-28