Kimi K3 Coding Tests Near the Frontier
PawelHuryn · x · 2026-07-17
The author compared Kimi K3 alongside Opus 4.8, GPT-5.6, and Grok 4.5 using the same 8-task coding benchmark.
Test Results
- Out of the 7 tasks K3 completed, it matched or exceeded the other three models in 6 of them.
- It caught 14 out of 21 planted bugs, beating Opus's 12, while the other models caught no more than 7.
- However, it also reported 2 false positives, making it the only model to do so.
Key Issues
- On launch night, 2 tasks returned 0 tokens directly—not wrong answers, but no output at all.
- After rerunning 13 hours later, 1 of those tasks achieved a perfect 14/14, indicating massive variability in model availability at different times.
- As of posting, OpenRouter is still rate-limiting, and Moonshot's native API is showing an engine overloaded warning.
Conclusion
- This is a 2.8 trillion parameter model, claiming to be the largest open-weight model in history.
- But the author's verdict is blunt: The weights are open, but the capacity is not. Currently, average users can't run it, and even Moonshot itself lacks the infrastructure to support it stably.
Related event: Kimi K3 Coding Test Nears Frontier Models but Lacks Usability(3 posts)→
More from coding & agent
- GPT-6 Astra beats Factorio with enemies in 44 in-game hours at ~$4,500 API cost — liminal_bardo · 2026-09-11
- 105 hidden bugs, 2 repos: DeepSeek V4.1 Flash fixes 24 at $1.80 vs Opus 5's 27 at $51.33 — ChartsJournalX · 2026-09-11
- Investment Analyst Asks How to Build a Claude-Based Diligence Agent Stack — Careless_Tie2286 · 2026-09-11
- Treating agents like 50 First Dates: a 3-layer context system so every conversation doesn't start from zero — evielync · 2026-09-11
- Running the Firefox MCP on Android via Termux, ngrok, and mcp-proxy — Nervous-Strain7544 · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11