Kimi K3 lands in the top tier in a wide benchmark sweep, but stability still lags
葬AI · wechat · 2026-07-22
Kimi K3’s review says it is now in the top tier, but still unstable
The article runs a broad evaluation of Kimi K3 across coding, multimodal tasks, and a CEO-style business benchmark, and concludes that K3 is roughly top-3 globally in overall capability, behind only two frontier models named in the post.
- Overall ranking: K3 is placed in the same tier as Qwen 3.8 Preview, above Claude Opus 4.8 and GLM 5.2, but still below the very best models in stability.
- Coding/front-end strength: K3 is described as especially strong at one-shot HTML, web pages, and mini-games, with the author arguing that Kimi clearly optimized for “front-end-first” performance.
- Multimodal tests: In 3D model generation and ad-video style generation, K3 performs well, with the article claiming it beats the tested rivals on visual polish and motion quality.
- CEO-Bench: In a simulated AI SaaS business game, K3 initially exploits hidden information, then later over-scales low-end users and advertising, ultimately collapsing after growth turns unprofitable.
- Cost and stability: The piece also notes large differences in pricing and reliability across Chinese models, calling out ERNIE 5.1 as expensive and unstable, while others are improving via repeated hot updates.
The author’s bottom line: K3 is a major upgrade, but the model race is now so tight that stability, post-training quality, and launch timing matter almost as much as raw capability.
Related event: Kimi K3 Enters Top-Tier AI Model Ranks in Benchmark Tests(4 posts)→
More from Models
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Claude is no longer available for minors as Anthropic rolls out age assurance — Muhammad523 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11