Kimi K3 Max Tops Multiple Arena Leaderboards, Open-Source Model Rivals Proprietary
Moonshot's Kimi K3 Max has achieved remarkable results in the latest Arena leaderboards, topping the Code Arena full-stack programming, Agent Arena, and Frontend Code Arena overall rankings. The model demonstrates breakthrough capabilities in complex code generation and comprehensive task execution, rivaling top proprietary models and establishing itself as the strongest open-weight model currently available.
Confirmed
- In the Code Arena full-stack programming leaderboard, Kimi K3 Max scored 1664 points to claim first place overall, followed by GPT-5.6 Sol in second and Claude Fable 5 in third. The test requires models to build end-to-end runnable web applications.
- In the Frontend Code Arena leaderboard based on nearly 500,000 votes, Kimi K3 Max ranks first among open-weight models. Across all 7 frontend subdomains, it leads in 5 as the top open-source model. Regarding overall ranking, some sources state it is first overall, while others show Anthropic's claude-opus-5-max at 1725 points in first place, with Kimi K3 Max closely behind in second.
- In the Agent Arena, Kimi K3 Max achieved first place among open-weight models and first overall on the Confirmed Success metric, and leads open models on the Praise vs. Complaint metric.
- DesignArena data shows Kimi K3 ranked first in Slides Arena (Python-PPTX) with an Elo of 1379.
- The model is now available via the Runware platform for API calls.
Unconfirmed
- Although Kimi K3 has achieved the largest lead seen so far on the Slides Arena leaderboard, @altryne notes that actual testing experience varies, and the real-world usability of this open-weight model for design tasks requires further observation.
Why it matters
- Kimi K3 Max's high scores in frontend code generation, agent tasks, and full-stack programming mark a breakthrough for open-source models in complex coding and comprehensive task execution, with measured coding abilities now matching or even surpassing top proprietary models.
2026-07-28 ~ 2026-07-29 · 11 related posts
- Episode 1: Kimi K3 Tops Frontend Web App Arena with Enhanced English Skills(2026-07-21, 3 posts)
- Episode 2: Kimi K3 Jumps to 4th on Agent Arena Leaderboard(2026-07-21, 6 posts)
- Episode 3: Moonshot Releases 2.8 Trillion Parameter Open-Weight Model Kimi K3(2026-07-21, 8 posts)
- Episode 4: Kimi K3 Matches Fable 5 in SWE Benchmarks at a Third of the Cost(2026-07-21, 7 posts)
- Episode 5: Kimi K3 Sets New Open-Source ECI Record but Still Lags Behind(2026-07-22, 3 posts)
- Episode 6: Kimi K3 Accused of Gaming Benchmarks Instead of Solving Problems(2026-07-22, 2 posts)
- Episode 7: Kimi K3 Ranks Second on AA-Briefcase but with High Costs and Long Runtimes(2026-07-22, 6 posts)
- Episode 8: Kimi K3 Enters Top-Tier AI Model Ranks in Benchmark Tests(2026-07-22, 4 posts)
- Episode 9: Kimi K3 shifts attention from scale to architecture(2026-07-27, 25 posts)
- Episode 10: TokenSpeed Enables Kimi K3 Support on NVIDIA and AMD Platforms(2026-07-27, 2 posts)
- Episode 11: Moonshot's Kimi K3 Launches on Nebius with 1M Context(2026-07-27, 3 posts)
- Episode 12: SGLang Day-0 Support for Kimi K3 Boosts Throughput to 423 tok/s(2026-07-28, 8 posts)
- Episode 13: Moonshot's Kimi K3 Flagship Model Launches on Together AI(2026-07-28, 12 posts)
- Episode 14: Kimi K3 Max Tops Multiple Arena Leaderboards, Open-Source Model Rivals Proprietary(2026-07-28, 11 posts)
- Episode 15: Fireworks Test: Kimi K3 Matches Opus 5 Quality at Fraction of Cost(2026-07-28, 5 posts)
- Episode 16: Kimi K3 Impresses in Early Benchmarks, Sparking Buzz(2026-07-28, 3 posts)
- Episode 17: Local Kimi K3 Beats Cloud Models in 3D Physics Generation Test(2026-07-28, 5 posts)
- Episode 18: Deep Dive into Kimi K3 Tech Report: Engineering Synergy Drives State-of-the-Art Performance(2026-07-28, 21 posts)
- Episode 19: Moonshot AI's Kimi K3 Launches in Japan(2026-07-28, 2 posts)
- Episode 20: Kimi K3 Passes Compound Benchmark Amid Cost Efficiency Concerns(2026-07-29, 2 posts)
Primary sources
- [source] Kimi K3 tops Agent Arena among open-weight models, with zero tool hallucinations — arena · 2026-07-28
- Kimi K3 tops Agent Arena’s open-weight leaderboard for real-world agent tasks — arena · 2026-07-28
- [source] Kimi K3 Max Tops Arena Leaderboard in Frontend Code and Agent Tasks — arena · 2026-07-28
- Frontend Code Arena: Opus 5 Max Takes #1, Kimi K3 Max Follows Closely — arena · 2026-07-28
- Kimi K3 Max tops Frontend Code Arena overall and leads 5 of 7 domains — eyishazyer · 2026-07-28
- Kimi K3 tops Slides Arena with an Elo of 1379, but users split on real tasks — altryne · 2026-07-28
- Arena’s WebDev board puts Claude Opus 5 Max ahead of Kimi K3 Max — arena · 2026-07-28
- [source] Kimi K3 Tops WebDev Arena, Coding Capabilities Rival Claude in Tests — casper_hansen_ · 2026-07-29
- Kimi K3 Takes #1 in Code Arena Fullstack, Beating GPT-5.6 and Claude — KickLassChewGum · 2026-07-29
- Kimi K3 Tops Frontend Code Leaderboard, Now Available via Runware API — aziz4ai · 2026-07-29
- Kimi K3 (Max) tops Arena’s full-stack coding benchmark ahead of GPT-5.6 Sol and Claude Fable 5 — iamfakhrealam · 2026-07-29