Kimi K3 Tests Near Frontier but Hard to Use
PawelHuryn · x · 2026-07-17
The author tested Kimi K3 alongside Opus 4.8, GPT-5.6, and Grok 4.5 using the same set of 8 tasks.
In the results, K3 performed on par or better in 6 out of 7 completed tasks, catching 14 out of 21 planted bugs, surpassing Opus's 12; however, it also produced 2 false positives, making it the only model to do so.
But the issues are also obvious:
- Right after release, 2 tasks directly returned 0 tokens
- Rerunning 13 hours later, one task achieved a perfect score, while the other still failed
- OpenRouter is still rate limiting, and Moonshot's own API shows "engine overloaded" errors
The author concludes that while K3's weights are open, the currently available inference capacity hasn't caught up yet.
Related event: Kimi K3 Coding Test Nears Frontier Models but Lacks Usability(3 posts)→
More from Models
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Claude is no longer available for minors as Anthropic rolls out age assurance — Muhammad523 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11
- Rumor claims Kimi faked performance by serving Claude; DeepSeek new model surprises in evals — realsohamparekh · 2026-09-11