Kimi K3 Architecture Analyzed: 3T Parameters with Sparser MoE, Compute Details Undisclosed
stochasticchasm · x · 2026-07-28
A developer analyzed the architecture of Moonshot's newly open-sourced Kimi K3 model. The model features 3T total parameters and adopts a sparser Mixture-of-Experts (MoE) configuration (16/896) compared to K2's 8/384.
The author notes the current sparsity is solid but could probably go even sparser, guessing that the bottleneck might be VRAM. Furthermore, they point out an interesting choice by the team: the technical reports for both Kimi K3 and Moonvit 2 do not disclose the training tokens or compute used anywhere.
Related event: Kimi K3 Architecture: 3T Parameters and Ultra-Sparse MoE(2 posts)→
More from Models
- Polymarket now prices a 42% chance of a new Claude Sonnet by next month — Polymarket · 2026-07-28
- Kimi K3 paper details SiTU-GLU and quantile balancing for 896-expert MoE — KyeGomezB · 2026-07-28
- vLLM Collaborates with DigitalOcean to Host Kimi K3 Model — vllm_project · 2026-07-28
- Kimi K3 lands on ChatLLM with U.S. hosting and an open-source fine-tune — bindureddy · 2026-07-28
- Kimi K3’s MoE routing may be driving higher expert-parallel communication costs — stochasticchasm · 2026-07-28
- Kimi K3 finds 16 new vulnerabilities and beats GLM-5.2 on an exploit benchmark — zephyr_z9 · 2026-07-28