Kimi K3 arrives as a 2.8T MoE with 104B active params and a 1M-token context
stochasticchasm · x · 2026-07-28
A reply to the Kimi K3 report notes that the tech report does not disclose training tokens or compute for either K3 or MoonVIT 2.
The attached abstract shows why the model is notable: Kimi K3 is a 2.8T-parameter MoE with 104B activated parameters, native vision support, and a 1M-token context window. The paper claims roughly a 2.5× scaling-efficiency improvement over Kimi K2, cites work on KDA, Stable LatentMoE, and long-horizon post-training, and says K3 reaches frontier-level results on long-context coding, agentic, knowledge, reasoning, and vision tasks.
It also says K3 still trails Claude Fable 5 and GPT-5.6 Sol overall, but beats other open and proprietary models in the authors’ evaluation suite. The full weights are released to support further research and adoption.
More from Models
- Kimi K3’s MoE routing may be driving higher expert-parallel communication costs — stochasticchasm · 2026-07-28
- Kimi K3 finds 16 new vulnerabilities and beats GLM-5.2 on an exploit benchmark — zephyr_z9 · 2026-07-28
- Kimi K3 weight shard appears as `model-00001-of-000096.safetensors` — ricklamers · 2026-07-28
- Microsoft launches MAI-Cyber-1-Flash and MDASH, claiming top CyberGym results at half the cost — satyanadella · 2026-07-28
- Claude Opus 5’s migration guide quietly changes years of prompting advice — AlexKim · 2026-07-28
- Post says an attention-heavy model has 104B active parameters — zephyr_z9 · 2026-07-28