SGLang Ecosystem Relay: Kimi K3 Hits 423 tok/s on Day-0 with Deep Optimizations
songhan_mit · x · 2026-07-30
The open-source inference framework SGLang achieved up to 423 tok/s on Kimi K3 on its day-0 release (measured on GSM8K), with native RL support ready.
This was powered by a massive relay race across the open-source ecosystem and tech giants. NVIDIA and AMD contributed serious kernel and hardware enablement, while KVCache pushed support for PD disaggregation and HiCache. Cloud providers like Modal, Baseten, and DigitalOcean engaged in co-development, testing, and compute provisioning.
SGLang deeply optimized K3's new architecture using fused KDA decode kernels, DP attention, DSpark, PD disagg, and KDA-aware prefix caching, enabling production-ready efficiency for the largest open-source model.
Related event: Moonshot Releases 2.8T Open-Source Model Kimi K3(3 posts)→
More from Infra
- TensorSharp Adds Multi-GPU Tensor Parallelism, Boosting Local GGUF Inference Speeds — fuzhongkai · 2026-07-30
- NVIDIA Open-Sources PyCuTe: Pure Python Layout Algebra for CUTLASS — asdf1234_0 · 2026-07-30
- UC Berkeley's K-search: Auto-Translating CUDA Kernel Optimizations to Apple's MLX — berkeley_ai · 2026-07-30
- Advantech Edge Device Powered by Nvidia Thor Runs RealSense GMSL Cameras — chrismatthieu · 2026-07-30
- Together Offers Lowest Price and Highest Cache Hit Rate for Kimi K3 on OpenRouter — zhyncs42 · 2026-07-30
- Samsung's Q2 Operating Profit Surges 1,800% to Record High Amid AI Chip Boom — Polymarket · 2026-07-30