Moonshot releases Kimi K3 report, claiming 2.5× scaling-efficiency gain over K2
cedric_chee · x · 2026-07-27
Moonshot has released the Kimi K3 technical report, claiming a 2.5× gain in scaling efficiency over Kimi K2.
The report attributes the improvement to a combination of architecture, data, and training changes. The attached table highlights the main structural differences, including a larger MoE stack, more activated parameters, more experts per token, a longer training context, and a shift to a hybrid KDA–MLA attention design.
Related event: Moonshot Releases Kimi K3 Report Claiming 2.5x Scaling Efficiency(3 posts)→
More from Models
- Kimi K3 launches on SGLang with 423 tok/s and 11 cloud partners — ying11231 · 2026-07-28
- Ollama Adds Kimi K3: 1M Context Window and Native Vision Support — ollama · 2026-07-27
- Kimi K3 reportedly improves training efficiency by 2.5× — zephyr_z9 · 2026-07-27
- NVIDIA distills Cosmos3 Super image-to-video to 4 steps with a 64B model — multimodalart · 2026-07-27
- Kimi K3 goes live on Modal with custom DFlash speculative decoding for lossless speedup — AAAzzam · 2026-07-27
- Kimi K3 with 2.8T parameters and 1M context now supported on vLLM — ricklamers · 2026-07-27