Kimi K3 Coding Tests Near the Frontier
PawelHuryn · x · 2026-07-17
The author compared Kimi K3 alongside Opus 4.8, GPT-5.6, and Grok 4.5 using the same 8-task coding benchmark.
Test Results
- Out of the 7 tasks K3 completed, it matched or exceeded the other three models in 6 of them.
- It caught 14 out of 21 planted bugs, beating Opus's 12, while the other models caught no more than 7.
- However, it also reported 2 false positives, making it the only model to do so.
Key Issues
- On launch night, 2 tasks returned 0 tokens directly—not wrong answers, but no output at all.
- After rerunning 13 hours later, 1 of those tasks achieved a perfect 14/14, indicating massive variability in model availability at different times.
- As of posting, OpenRouter is still rate-limiting, and Moonshot's native API is showing an engine overloaded warning.
Conclusion
- This is a 2.8 trillion parameter model, claiming to be the largest open-weight model in history.
- But the author's verdict is blunt: The weights are open, but the capacity is not. Currently, average users can't run it, and even Moonshot itself lacks the infrastructure to support it stably.
Related event: Kimi K3 Coding Test Nears Frontier Models but Lacks Usability(3 posts)→
More from coding & agent
- Codex helps build Valdiluce, an open-world game with climbing, gliding and gondolas — Dimillian · 2026-07-22
- HeyGen adds a media-sourcing skill for coding agents with 75k images and 10k tracks — HeyGen · 2026-07-22
- Agent search bottlenecks are now about variance, not raw latency — rohanpaul_ai · 2026-07-22
- LangSmith adds tracing for Pipecat, LiveKit, OpenAI Realtime, and Gemini Live — LangChain · 2026-07-22
- An MCP server signs every AI agent tool call into a verifiable Merkle chain — Funky_Chicken_22 · 2026-07-22
- Annotated transcript of a Claude Code team interview is now available — trq212 · 2026-07-22