Moonshot releases Kimi K3, a 2.8T multimodal MoE with 1M-token context
月之暗面 Kimi · wechat · 2026-07-27
Moonshot released Kimi K3 and opened the model weights, plus the key infrastructure behind training it.
- Model: a 2.8T-parameter MoE model with native vision understanding and a 1M-token context window.
- Efficiency: Moonshot says the model is about 3× the size of Kimi K2.5, but training efficiency improved by 2.5× thanks to a set of scaling and systems optimizations.
- Technical report: covers KDA+AttnRes, StableLatentMoE, MoonViT-V2, and post-training / evaluation across general reasoning, agent tasks, and coding agents.
- Open-sourced infra: MoonEP (MoE communication), FlashKDA (high-performance kernel), and AgentEnv (sandbox for large-scale agent RL workflows).
- Moonshot says the weights are free to download and deploy for internal R&D and end-user products, subject to the published terms.
Related event: Moonshot Open-Sources Kimi K3: 2.8T Params and $20M Commercial Threshold(148 posts)→
More from Infra
- Open-source proxy stops agent loops, caps budget, and blocks risky tool calls — Electrical-War-549 · 2026-07-28
- FT screenshots show AI chip optimism and a sharp sell-off in the same day — nathanbenaich · 2026-07-28
- Kimi-K3 hosts appear locked into identical pricing as providers compete on latency — TheZachMueller · 2026-07-28
- Alphabet’s estimated 12-month capex jumps to $250B as the AI race heats up — FinanceYF5 · 2026-07-28
- Samsung chip workers are leaving for SK Hynix as $476,000 HBM bonuses lure talent — nordicinst · 2026-07-28
- Hydra routes local tasks to the cheapest model that clears a confidence threshold — jhaankit373 · 2026-07-28