Kimi K3 Beats DeepSeek-V4 in Code Arena, But DeepSeek is 14x Cheaper
rohanpaul_ai · x · 2026-08-14
In the latest early AutoEval for Code Arena: WebDev, Kimi K3 (Max) leads open-source models with 1674 points, followed closely by DeepSeek-V4-Pro (Max) at 1607 points.
Despite Kimi's 67-point lead, its API pricing is significantly higher. For a 1M input + 1M output workload, Kimi costs around $18, whereas DeepSeek costs only $1.30, making it nearly 14x more cost-effective.
The analysis notes that the Arena score gap doesn't translate linearly to code quality. Kimi's premium is only justified if it materially reduces failed edits and human reviews in actual development environments.
More from Models
- Anthropic's Suspicious Silence Hints at Upcoming 'Claudette' Drop — beffjezos · 2026-08-14
- DeepSeek V4-Pro slows to crawl on long projects; reloading reveals it's further ahead — teortaxesTex · 2026-08-14
- Google Focuses on Smaller, Faster AI Models Leveraging Massive Search Scale — haider1 · 2026-08-14
- Notion Launches Knowledge Board: Evaluating LLMs on Real-World Traffic Instead of Benchmarks — ivanhzhao · 2026-08-14
- DeepSeek V4 Pro Hits Baseten APIs: 1.7T Parameters, MIT License — baseten · 2026-08-14
- Grok Blocks Intimate Likeness Edits, Proving Musk's 'Legal = Allowed' Wrong — firasd · 2026-08-14