Moonshot releases Kimi K3 report, claiming 2.5× scaling-efficiency gain over K2

cedric_chee · x · 2026-07-27

Moonshot has released the Kimi K3 technical report, claiming a 2.5× gain in scaling efficiency over Kimi K2.

The report attributes the improvement to a combination of architecture, data, and training changes. The attached table highlights the main structural differences, including a larger MoE stack, more activated parameters, more experts per token, a longer training context, and a shift to a hybrid KDA–MLA attention design.

Related event: Moonshot Releases Kimi K3 Report Claiming 2.5x Scaling Efficiency(3 posts)→

Original post →

More from Models

Models channel →