SGLang Ecosystem Relay: Kimi K3 Hits 423 tok/s on Day-0 with Deep Optimizations

songhan_mit · x · 2026-07-30

The open-source inference framework SGLang achieved up to 423 tok/s on Kimi K3 on its day-0 release (measured on GSM8K), with native RL support ready.

This was powered by a massive relay race across the open-source ecosystem and tech giants. NVIDIA and AMD contributed serious kernel and hardware enablement, while KVCache pushed support for PD disaggregation and HiCache. Cloud providers like Modal, Baseten, and DigitalOcean engaged in co-development, testing, and compute provisioning.

SGLang deeply optimized K3's new architecture using fused KDA decode kernels, DP attention, DSpark, PD disagg, and KDA-aware prefix caching, enabling production-ready efficiency for the largest open-source model.

Related event: Moonshot Releases 2.8T Open-Source Model Kimi K3(3 posts)→

Original post →

More from Infra

Infra channel →