Moonshot releases Kimi K3, a 2.8T MoE model with 1M context and 423 tok/s serving
ricklamers · x · 2026-07-27
Moonshot released Kimi K3, describing it as its most capable model yet:
- A 2.8T MoE model with native visual understanding.
- A 1M-token context window.
- The company says the new architecture delivers 2.5x more intelligence per unit of compute, not just more parameters.
- Alongside the model, Moonshot is opening more of the stack: high-performance attention kernels, an MoE communication library, and infrastructure for agent environments at scale.
The quoted SGLang post adds that K3 reaches 423 tok/s on gsm8k day one, with RL support ready, and that the serving stack uses fused KDA decode kernels, DP attention, DSpark, PD disagg, and KDA-aware prefix caching. The demo video was reportedly generated by Kimi K3 itself.
Related event: Moonshot releases Kimi K3 open weights amid license debate(155 posts)→
More from Infra
- Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference — togethercompute · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23
- Qwen 27B runs 24hr unattended on one RTX5090, builds full Postgres-SpringBoot-React spreadsheet app — anglepoiselife · 2026-09-23
- OpenRoboto Shift launches: decentralized egocentric video data network for robot brains — markjeffrey · 2026-09-23
- Engineer describes designing digital circuits that recycle most of their energy — MikePFrank · 2026-09-23