Kimi K3 Bottoms Out in Real-World Repair Tests
ChrisUniverse · x · 2026-07-18
A quoted post points out that the narrative surrounding Kimi K3's strong coding capabilities is inconsistent with their own test results.
Key information provided includes:
- K3 ranks 1st in Arena Frontend Code with a score of 1679
- Ranks 3rd in Artificial Analysis with an Intelligence Index of 57
- However, on their coding-agent repair harness, K3 ranked last out of 7 models
- It achieved 53/67 successful repairs, costing about $0.186 per successful fix, taking an average of 702 seconds
- In contrast, Sol achieved 100% (70/70), and Grok hit 99% with an average time of only 46 seconds
The author's point is that the external impression of K3 being "crushingly good" might starkly contrast with its performance under specific real-world workloads.
Related event: Kimi K3 Evaluations Show Polarized Results and Harness Sensitivity(5 posts)→
More from Models
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11
- GPT-5.6 writes well but is instantly forgettable, user complains — BasedRaddka · 2026-09-11
- Opus Refuses Protein Research Codebase Over 'Safety' Concerns, Dev Considers Rolling His Own — josephdviviano · 2026-09-11
- User Hails Unconfirmed 'DeepSeek 4.1 Flash' as an Inflection Point in LLMs — himanshustwts · 2026-09-11