Moonshot's Kimi K3 Tops Frontend Code Arena, Nearing Fable 5 in Coding at a Third of the Cost
From July 18 to 19, Moonshot's new model Kimi K3 put its coding ability at the center of community attention. Arena officially announced that Kimi-K3 scored 1,679 on the Frontend Code Arena and rose to the top, making a Chinese model lead that US-dominated leaderboard for the first time; it was also reported to take first place on the WebDev human-preference chart. A wave of analyses followed around DeepSWE, Artificial Analysis and other benchmarks, converging on one consensus: Kimi K3's coding ability now approaches top-tier closed models, at roughly a third of the price.
Key Results
Beyond the Arena leaderboard, reposts claimed Kimi K3 also surpassed Claude Fable 5 and other closed models on terminal and long-horizon coding tasks closer to real product development. In Artificial Analysis's coding Agent index it scored 57, tied for fifth with GPT-5.6 Terra and ahead of Opus 4.8. A multi-sampling comparison reshared by zephyr_z9 showed Kimi K3's pass@1 at 68.5, slightly below Fable 5's 69.9, but its pass@2 rising to 82.0.
Value and Language-Level Performance
rohanpaul_ai cited a comparison chart showing that on DeepSWE rollout tasks, measured per $100 of cost, Kimi completes 14.7 tasks versus Fable 5's 5.3. zhyncs42 relayed that Kimi K3 can cut frontier-model inference cost by about 3x and will be natively available on Together Compute from July 27. ZainHasan6 broke it down by language: on Rust it comes extremely close to top-ranked Fable and even beats GPT 5.6 Sol, while on Go it defeats Fable 79 to 71.
Failure Modes and Consistency
ZainHasan6 reported per-task consistency stats: the correlation between Kimi K3 and Fable reaches 0.72, the highest cross-vendor similarity he has seen; both pass 96 tasks, with Kimi-only 5 and Fable-only 15. Comparing failure distributions, he judged the two models' "failure fingerprints" to be nearly identical, with about 65% of failures being near misses and a stronger tendency toward conservative failures rather than hallucinatory errors. Taken together, this round of discussion positions Kimi K3 as a new contender that balances benchmark scores with cost in coding scenarios.
2026-07-18 ~ 2026-07-19 · 12 related posts
- [source] Kimi-K3 Tops Frontend Code Arena — arena · 2026-07-18
- Kimi K3 Ties for 5th in Coding, Beats GPT-5.6 in Value — scaling01 · 2026-07-18
- [source] Kimi K3 and Fable Show High Evaluation Consistency — ZainHasan6 · 2026-07-18
- Kimi K3 and Fable Share Strikingly Similar Failure Modes — ZainHasan6 · 2026-07-18
- Kimi K3 vs. Fable 5 Coding Evaluation Comparison — ZainHasan6 · 2026-07-18
- Kimi K3 vs Fable 5 Evaluation — ZainHasan6 · 2026-07-18
- Kimi K3 Claims to Cut Inference Costs to a Third — zhyncs42 · 2026-07-18
- Kimi K3 Leads Multi-Sampling Benchmarks — zephyr_z9 · 2026-07-18
- Kimi K3 Shines in Coding: Matches Fable in Rust — ZainHasan6 · 2026-07-18
- Kimi K3 Tops WebDev Leaderboard — ylecun · 2026-07-19
- [source] Kimi K3 Offers Better Cost-Performance for Coding Tasks — rohanpaul_ai · 2026-07-19
- Kimi K3 Tops Code Arena Leaderboard — koltregaskes · 2026-07-19