Kimi K3 arrives as a 2.8T MoE with 104B active params and a 1M-token context
stochasticchasm · x · 2026-07-28
A reply to the Kimi K3 report notes that the tech report does not disclose training tokens or compute for either K3 or MoonVIT 2.
The attached abstract shows why the model is notable: Kimi K3 is a 2.8T-parameter MoE with 104B activated parameters, native vision support, and a 1M-token context window. The paper claims roughly a 2.5× scaling-efficiency improvement over Kimi K2, cites work on KDA, Stable LatentMoE, and long-horizon post-training, and says K3 reaches frontier-level results on long-context coding, agentic, knowledge, reasoning, and vision tasks.
It also says K3 still trails Claude Fable 5 and GPT-5.6 Sol overall, but beats other open and proprietary models in the authors’ evaluation suite. The full weights are released to support further research and adoption.
Related event: Moonshot releases Kimi K3 open weights amid license debate(155 posts)→
More from Models
- theo builds his own visualizer for today's agent models, showing how cheap Luna really is — ivan_bezdomny · 2026-09-23
- Why ChatGPT Still Wins: One User's Split Between Muse, Claude and Codex — mobileraj · 2026-09-23
- Muse reportedly offers 4B tokens/week for ~$100/month, sparking industry price-disruption talk — NewYak4281 · 2026-09-23
- GPT-6 Sol and Luna appear in OpenAI docs, alongside guidance on reasoning effort — cedric_chee · 2026-09-23
- GPT-6 tested on LIBERO robot task: turns on stove, fails to grasp moka pot — YuXiang_IRVL · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23