SGLang says Kimi K3 hits 423 tok/s on day 0 with fused KDA kernels and prefix caching
ying11231 · x · 2026-07-29
- SGLang reports 423 tok/s on Kimi K3 at day 0, measured on GSM8K, and says RL support is already ready in Miles.
- The team says K3’s new architecture was natively implemented and deeply optimized with fused KDA decode kernels, DP attention, DSpark, PD disagg, and KDA-aware prefix caching.
- They also claim to have passed Kimi Vendor Verifier and say the stack is ready for production.
- The post credits collaboration with Kimi Moonshot, NVIDIA, AMD, KVCache AI, Modal, Baseten, and several serving partners.
- Blog, cookbook, and benchmarks were promised in the comments, and the demo video was described as being generated by Kimi K3 itself.
Related event: SGLang Launches Day-0 Support for Kimi K3, Hitting 423 tok/s(8 posts)→
More from Infra
- YC startup hwintelligence launches Wave, a waveform debugger with a verification agent — ycombinator · 2026-07-29
- Broadcom says AI could cut exploit windows from weeks to hours — therealdanvega · 2026-07-29
- camelAI moved its agent from VMs to a Cloudflare Durable Object — irvinebroque · 2026-07-29
- LLM Inference Handbook collects deployment, GPU, and optimization guidance for production teams — carrycooldude · 2026-07-29
- MLCommons launches MLPerf Endpoints v0.7 for AI inference benchmarking — TheKanter · 2026-07-29
- Solo project compares cloud providers and generates Terraform — Character-Ring5785 · 2026-07-29