Kimi K3 Tests Near Frontier but Hard to Use

PawelHuryn · x · 2026-07-17

The author tested Kimi K3 alongside Opus 4.8, GPT-5.6, and Grok 4.5 using the same set of 8 tasks.

In the results, K3 performed on par or better in 6 out of 7 completed tasks, catching 14 out of 21 planted bugs, surpassing Opus's 12; however, it also produced 2 false positives, making it the only model to do so.

But the issues are also obvious:

The author concludes that while K3's weights are open, the currently available inference capacity hasn't caught up yet.

Related event: Kimi K3 Coding Test Nears Frontier Models but Lacks Usability(3 posts)→

Original post →

More from Models

Models channel →