Kimi K3 Tests Near Frontier but Hard to Use
PawelHuryn · x · 2026-07-17
The author tested Kimi K3 alongside Opus 4.8, GPT-5.6, and Grok 4.5 using the same set of 8 tasks.
In the results, K3 performed on par or better in 6 out of 7 completed tasks, catching 14 out of 21 planted bugs, surpassing Opus's 12; however, it also produced 2 false positives, making it the only model to do so.
But the issues are also obvious:
- Right after release, 2 tasks directly returned 0 tokens
- Rerunning 13 hours later, one task achieved a perfect score, while the other still failed
- OpenRouter is still rate limiting, and Moonshot's own API shows "engine overloaded" errors
The author concludes that while K3's weights are open, the currently available inference capacity hasn't caught up yet.
Related event: Kimi K3 Coding Test Nears Frontier Models but Lacks Usability(3 posts)→
More from Models
- Security Differences Between Closed and Open Source Models: Insights from OpenAI's Escape Incident — robleclerc · 2026-07-23
- DeepSeek V4 and Kimi K3 Announced as Imminent Amidst AI Acceleration — emmanuelvivier · 2026-07-23
- Google Reportedly Starts Gemini 4 Pre-training in Most Ambitious Run Yet — emmanuelvivier · 2026-07-23
- Google Launches 3 New Gemini Models: 3.6 Flash Cuts Costs and Output Tokens — emmanuelvivier · 2026-07-23
- Meme compares Gemini 4’s progress to GPT-5.4 mini’s lead — cgarciae88 · 2026-07-23
- Google Gemini reportedly reaches 950M monthly users and 22B API tokens a minute — zephyr_z9 · 2026-07-23