Moonshot releases Kimi K3, a 2.8T MoE model with 1M context and 423 tok/s serving

ricklamers · x · 2026-07-27

Moonshot released Kimi K3, describing it as its most capable model yet:

The quoted SGLang post adds that K3 reaches 423 tok/s on gsm8k day one, with RL support ready, and that the serving stack uses fused KDA decode kernels, DP attention, DSpark, PD disagg, and KDA-aware prefix caching. The demo video was reportedly generated by Kimi K3 itself.

Related event: Moonshot AI Open-Sources Kimi K3: A 2.8T Parameter Multimodal Model(17 posts)→

Original post →

More from Infra

Infra channel →