WebDev AI Arena: Claude Dominates Leaderboard, Kimi Takes Second
arena · x · 2026-08-01
LMSYS has updated its WebDev AI Leaderboard, ranking frontier models on front-end web development tasks based on over 500,000 community votes. The evaluation includes agentic coding workflows requiring multi-step reasoning and tool use.
Top Contenders:
- claude-opus-5-max (Anthropic) ranks 1st with a score of 1704, priced at $5/$25 per million tokens.
- kimi-k3-max (Moonshot) takes an impressive 2nd place at 1675 points, offering significant cost efficiency at $3/$15.
- claude-opus-5-high (1664) and claude-fable-5 (1630) follow, giving Anthropic a clean sweep of the top four.
Other Major Models:
- OpenAI's gpt-5.6-sol-xhigh sits at 5th (1621).
- Among open-source models, Z.ai's glm-5.2-max (1587) and DeepSeek's deepseek-v4-flash-high (1586) break into the top tier, with the latter featuring an ultra-low price of $0.14/$0.28.
Related event: WebDev AI Model Arena Updates: Claude Tops, Kimi Ranks Second(2 posts)→
More from Models
- Lamenting Claude Haiku 3.5: Developers Urge AI Companies to Stop Deprecating Old Models — repligate · 2026-08-01
- DeepSeek Matches Claude Sonnet in Agentic Loops at 600% Lower Cost — bindureddy · 2026-08-01
- Kimi K3 Hits Record 172 Tokens/sec in Inference Speed — AccBalanced · 2026-08-01
- Kimi K3 Hits OpenRouter: 2.8T Parameters, 1M Context Length — AccBalanced · 2026-08-01
- DeepSeek Runs Locally on Workstations, Making Open-Weight AI Bans Impossible — pstAsiatech · 2026-08-01
- Claude Code Costs 3.7x More Than Open-Source Agents in Task Benchmark — Teknium · 2026-08-01