Moonshot releases Kimi K3, a 2.8T MoE model with 1M context and 423 tok/s serving
ricklamers · x · 2026-07-27
Moonshot released Kimi K3, describing it as its most capable model yet:
- A 2.8T MoE model with native visual understanding.
- A 1M-token context window.
- The company says the new architecture delivers 2.5x more intelligence per unit of compute, not just more parameters.
- Alongside the model, Moonshot is opening more of the stack: high-performance attention kernels, an MoE communication library, and infrastructure for agent environments at scale.
The quoted SGLang post adds that K3 reaches 423 tok/s on gsm8k day one, with RL support ready, and that the serving stack uses fused KDA decode kernels, DP attention, DSpark, PD disagg, and KDA-aware prefix caching. The demo video was reportedly generated by Kimi K3 itself.
Related event: Moonshot AI Open-Sources Kimi K3: A 2.8T Parameter Multimodal Model(17 posts)→
More from Infra
- Moonshot open-sources MoonEP as open models vs closed labs debate intensifies — KyeGomezB · 2026-07-27
- Kimi K3’s 2.5x scaling-law gain draws praise for training efficiency — andrew_n_carr · 2026-07-27
- Kimi K3 goes live on Nebius with 1M-token context and a 57 AA score — teortaxesTex · 2026-07-27
- AI agent finds a longstanding Bun Node-compat bug in `child_process.spawn` — steipete · 2026-07-27
- llama.cpp adds support for Nanbeige4.2 in pull request 25994 — pmttyji · 2026-07-27
- Open-sourced Kimi K3 speculator lifts single-stream throughput from 118 to 370 tok/s — vllm_project · 2026-07-27