Kimi K3 Ranks 2nd in Coding Test, Distillation Praised
teortaxesTex · x · 2026-07-19
ValsAI's internal benchmark, Vibe Code Bench, shows the Kimi K3 model ranking second overall with a score of 85.0%. The test primarily evaluates a model's ability to create web applications from scratch.
Commenter teortaxesTex expressed amazement, praising the Kimi team's excellent work on model distillation. They noted it is even closer to the Fable model's level than Anthropic's Sonnet 3.5, implying that US companies' legal compliance concerns might be limiting their models' performance ceilings.
More from Models
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11
- Terminal Bench v4: GLM-5.3 Leads at 41.9%, Kimi-K3 Underwhelms at 12.6% — Ok_Warning2146 · 2026-09-11
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11