Kimi K3 triples parameters and doubles active experts in a new MoE design
teortaxesTex · x · 2026-07-27
A technical comparison image for Kimi K3 highlights several major architecture changes versus K2:
- layers increased from 61 to 93
- total parameters rose from 1.04T to 2.78T
- activated parameters jumped from 32.6B to 104.2B
- routed experts increased from 384 to 896, while active experts per token doubled from 8 to 16
- context length expanded from 128K to 1M
- the model also adds a ViT component
The post argues that K3's latent bottleneck is not reducing routed traffic, but reallocating capacity to twice as many active experts, with per-token expert dispatch volume remaining unchanged from K2.
Related event: Moonshot releases Kimi K3 open weights amid license debate(155 posts)→
More from Models
- French prize-winning novel suspected of AI: $1,000 challenge over detector results — Afinetheorem · 2026-09-23
- Third-party test: Claude Opus 5.5 renders finer 3D scenes but costs 13x more than GPT-6 Sol — testingcatalog · 2026-09-23
- GPT-6 Sol priced at half of Opus 5.5 as Sol and Luna go 'dirt cheap' — ZeroStateReflex · 2026-09-23
- Tester claims Claude Opus 5.5 has the best visual design output of any model tested — burny_tech · 2026-09-23
- Meta's Alexandr Wang reveals muse has been in the works since at least Sept 2025 — adrianscottcom · 2026-09-23
- GPT-6 Sol Codex system prompt leaked: over 294,000 characters dumped on GitHub — gaganghotra_ · 2026-09-23