SGLang Day-0 Support for Kimi K3 Boosts Throughput to 423 tok/s
The SGLang team announced Day-0 support for the latest open-source model, Kimi K3. Through deep optimization for K3's new architecture and speculative decoding, this 2.8T parameter model with a 1M context window saw its batch-1 decode throughput on the GSM8K benchmark surge from about 113 tok/s to 423 tok/s, with reinforcement learning (RL) support already ready. This proves that system-level software optimizations can significantly break through the decoding bottlenecks of extremely large models.
Confirmed
- Model Scale and Capabilities: Kimi K3 is a 2.8T parameter open-source model supporting a 1M context window.
- Inference Performance and Optimization: K3 achieved an inference throughput of 423 tok/s on the GSM8K benchmark. SGLang made a series of adaptations for its hybrid KDA-MLA architecture, including fused KDA decoding kernels, DP attention, PD separation, KDA-aware prefix caching, unified memory, ReplaySSM, and chunked PP and decode CP. Meanwhile, the Radixark team used SpecForge to train a DSpark speculator draft model, successfully boosting K3's batch-1 decode throughput from approximately 113 tok/s to 423 tok/s. SGLang's officially released demo video was generated by K3 running on this same serving stack.
- Cross-Platform Deployment: K3 has been successfully run on AMD MI350X via SGLang, achieving 327 tok/s with four-way concurrency. The deployment process was described as almost out-of-the-box, with credits given to the AMD, SGLang, and Moonshot teams.
Why it matters
As an extremely large open-source model, K3's ability to achieve high inference throughput at launch proves that system-level software optimizations (like SGLang) and speculative decoding technologies (like DSpark) can drastically overcome the decoding bottlenecks of massive models. Furthermore, its smooth operation on AMD hardware provides developers with a viable computing alternative to Nvidia.
2026-07-28 ~ 2026-07-29 · 8 related posts
- Episode 1: Kimi K3 Tops Frontend Web App Arena with Enhanced English Skills(2026-07-21, 3 posts)
- Episode 2: Kimi K3 Jumps to 4th on Agent Arena Leaderboard(2026-07-21, 6 posts)
- Episode 3: Moonshot Releases 2.8 Trillion Parameter Open-Weight Model Kimi K3(2026-07-21, 8 posts)
- Episode 4: Kimi K3 Matches Fable 5 in SWE Benchmarks at a Third of the Cost(2026-07-21, 7 posts)
- Episode 5: Kimi K3 Sets New Open-Source ECI Record but Still Lags Behind(2026-07-22, 3 posts)
- Episode 6: Kimi K3 Accused of Gaming Benchmarks Instead of Solving Problems(2026-07-22, 2 posts)
- Episode 7: Kimi K3 Ranks Second on AA-Briefcase but with High Costs and Long Runtimes(2026-07-22, 6 posts)
- Episode 8: Kimi K3 Enters Top-Tier AI Model Ranks in Benchmark Tests(2026-07-22, 4 posts)
- Episode 9: Kimi K3 shifts attention from scale to architecture(2026-07-27, 25 posts)
- Episode 10: TokenSpeed Enables Kimi K3 Support on NVIDIA and AMD Platforms(2026-07-27, 2 posts)
- Episode 11: Moonshot's Kimi K3 Launches on Nebius with 1M Context(2026-07-27, 3 posts)
- Episode 12: SGLang Day-0 Support for Kimi K3 Boosts Throughput to 423 tok/s(2026-07-28, 8 posts)
- Episode 13: Moonshot's Kimi K3 Flagship Model Launches on Together AI(2026-07-28, 12 posts)
- Episode 14: Kimi K3 Max Tops Multiple Arena Leaderboards, Open-Source Model Rivals Proprietary(2026-07-28, 11 posts)
- Episode 15: Fireworks Test: Kimi K3 Matches Opus 5 Quality at Fraction of Cost(2026-07-28, 5 posts)
- Episode 16: Kimi K3 Impresses in Early Benchmarks, Sparking Buzz(2026-07-28, 3 posts)
- Episode 17: Local Kimi K3 Beats Cloud Models in 3D Physics Generation Test(2026-07-28, 5 posts)
- Episode 18: Deep Dive into Kimi K3 Tech Report: Engineering Synergy Drives State-of-the-Art Performance(2026-07-28, 21 posts)
- Episode 19: Moonshot AI's Kimi K3 Launches in Japan(2026-07-28, 2 posts)
- Episode 20: Kimi K3 Passes Compound Benchmark Amid Cost Efficiency Concerns(2026-07-29, 2 posts)
Primary sources
- Kimi K3 launches on SGLang with 423 tok/s and 11 cloud partners — ying11231 · 2026-07-28
- SGLang adds day-one Kimi K3 support and reports 423 tokens/s on GSM8K — BanghuaZ · 2026-07-28
- [source] Kimi K3 serving stack reaches 423 tok/s after DSpark draft-model tuning — ying11231 · 2026-07-28
- SGLang Day-0 Support for Kimi K3 Hits 423 tok/s on GSM8K — ying11231 · 2026-07-28
- [source] Kimi K3 runs on AMD MI350X with SGLang and hits 327 tok/s across four requests — burny_tech · 2026-07-28
- [source] SGLang adapts to Kimi K3’s hybrid architecture with memory and kernel optimizations — SonglinYang4 · 2026-07-28