Kimi K3 triples parameters and doubles active experts in a new MoE design
teortaxesTex · x · 2026-07-27
A technical comparison image for Kimi K3 highlights several major architecture changes versus K2:
- layers increased from 61 to 93
- total parameters rose from 1.04T to 2.78T
- activated parameters jumped from 32.6B to 104.2B
- routed experts increased from 384 to 896, while active experts per token doubled from 8 to 16
- context length expanded from 128K to 1M
- the model also adds a ViT component
The post argues that K3's latent bottleneck is not reducing routed traffic, but reallocating capacity to twice as many active experts, with per-token expert dispatch volume remaining unchanged from K2.
More from Models
- Kimi K3 launches on SGLang with 423 tok/s and 11 cloud partners — ying11231 · 2026-07-28
- Ollama Adds Kimi K3: 1M Context Window and Native Vision Support — ollama · 2026-07-27
- Kimi K3 reportedly improves training efficiency by 2.5× — zephyr_z9 · 2026-07-27
- NVIDIA distills Cosmos3 Super image-to-video to 4 steps with a 64B model — multimodalart · 2026-07-27
- Kimi K3 goes live on Modal with custom DFlash speculative decoding for lossless speedup — AAAzzam · 2026-07-27
- Kimi K3 with 2.8T parameters and 1M context now supported on vLLM — ricklamers · 2026-07-27