Moonshot releases Kimi K3, a 2.8T multimodal MoE with 1M-token context
月之暗面 Kimi · wechat · 2026-07-27
Moonshot released Kimi K3 and opened the model weights, plus the key infrastructure behind training it.
- Model: a 2.8T-parameter MoE model with native vision understanding and a 1M-token context window.
- Efficiency: Moonshot says the model is about 3× the size of Kimi K2.5, but training efficiency improved by 2.5× thanks to a set of scaling and systems optimizations.
- Technical report: covers KDA+AttnRes, StableLatentMoE, MoonViT-V2, and post-training / evaluation across general reasoning, agent tasks, and coding agents.
- Open-sourced infra: MoonEP (MoE communication), FlashKDA (high-performance kernel), and AgentEnv (sandbox for large-scale agent RL workflows).
- Moonshot says the weights are free to download and deploy for internal R&D and end-user products, subject to the published terms.
Related event: Moonshot releases Kimi K3 open weights amid license debate(155 posts)→
More from Infra
- 12 KV Cache Reduction Techniques Every AI Engineer Should Understand, Explained — blaizedsouza · 2026-09-11
- The shadow GPU capacity market is formalizing, with Meta selling excess compute to outside buyers — DavidLinthicum · 2026-09-11
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- 80% of the DIY LLM inference hype posters have already quit — it's brutally hard systems work — abhijithneil · 2026-09-11
- Hugging Face's Ultra Scale Playbook: a free book on training LLMs on GPU clusters — mdancho84 · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11