SGLang adds day-one Kimi K3 support and reports 423 tokens/s on GSM8K
BanghuaZ · x · 2026-07-28
SGLang announced day-0 support for Kimi K3, claiming a speed of 423 tokens/s on GSM8K and saying RL support is ready.
- The post describes K3 as a 2.8T-parameter open model with a 1M context window.
- SGLang says it natively implemented and deeply optimized K3’s architecture with fused KDA decode kernels, DP attention, DSpark, PD disaggregation, and KDA-aware prefix caching.
- It reports production readiness after passing the Kimi Vendor Verifier.
- The stack is said to run across NVIDIA GB300/B300/B200/H200/H20, AMD MI350X/MI355X, and more.
- The launch video was also generated by K3 itself and turned into a playable mini game.
Related event: Kimi K3 Hits 423 tok/s with SGLang Integration(4 posts)→
More from Infra
- SSI says NVIDIA investment will help it 10x compute in the next 12 months — vitaliychiley · 2026-07-28
- Macrocosmos starts a permissionless 16B model training run across three continents — markjeffrey · 2026-07-28
- Optimizing K8s Resource Requests Yields 9x Speedup for Whisper Workloads — anacondainc · 2026-07-28
- Celestica’s AI server margins fall from 12.8% to 10.8% as the economics come into focus — tengyanAI · 2026-07-28
- Linear-attention hybrids may need finer caching for long prompts and workflows — stochasticchasm · 2026-07-28
- A research-agent ranking of 8 stock-data MCP servers puts Equibles first — DanielAPO · 2026-07-28