SGLang says it can serve Kimi K3 at 423 tok/s with day-0 production support
vwxyzjn · x · 2026-07-28
SGLang says it can serve Kimi K3 at day-0 speed, measuring 423 tokens/s on GSM8K, and says RL support is already ready in Miles.
The post explains that SGLang natively implemented and optimized K3’s new architecture with fused KDA decode kernels, DP attention, DSpark, PD disaggregation, and KDA-aware prefix caching. It says the model passed Kimi Vendor Verifier and is ready for production, with benchmarks, a blog post, and a cookbook linked in the comments.
Related event: SGLang Day-0 Support for Kimi K3 Boosts Throughput to 423 tok/s(8 posts)→
More from Infra
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- 80% of the DIY LLM inference hype posters have already quit — it's brutally hard systems work — abhijithneil · 2026-09-11
- Hugging Face's Ultra Scale Playbook: a free book on training LLMs on GPU clusters — mdancho84 · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11