SGLang Day-0 Support for Kimi K3 Boosts Throughput to 423 tok/s

The SGLang team announced Day-0 support for the latest open-source model, Kimi K3. Through deep optimization for K3's new architecture and speculative decoding, this 2.8T parameter model with a 1M context window saw its batch-1 decode throughput on the GSM8K benchmark surge from about 113 tok/s to 423 tok/s, with reinforcement learning (RL) support already ready. This proves that system-level software optimizations can significantly break through the decoding bottlenecks of extremely large models.

Confirmed

Why it matters

As an extremely large open-source model, K3's ability to achieve high inference throughput at launch proves that system-level software optimizations (like SGLang) and speculative decoding technologies (like DSpark) can drastically overcome the decoding bottlenecks of massive models. Furthermore, its smooth operation on AMD hardware provides developers with a viable computing alternative to Nvidia.

2026-07-28 ~ 2026-07-29 · 8 related posts

Full story(20 episodes)→

Primary sources

2 near-duplicate retellings: vwxyzjn · ying11231