Kimi K3 serving stack reaches 423 tok/s after DSpark draft-model tuning
ying11231 · x · 2026-07-28
Radixark says they trained a DSpark speculator draft model for Kimi K3 with SpecForge, raising batch-1 decode throughput from about 113 tok/s to 423 tok/s.
Reported gains
- +68% throughput at bs=256 on chat vs. verify-all, lossless
- Accept length of 5.67 on GSM8K and 5.36 on HumanEval
The quoted SGLang update says K3’s fast serving comes from native implementation and optimization of the new architecture, including fused KDA decode kernels, DP attention, DSpark, PD disagg, and KDA-aware prefix caching. It also says the stack has passed Kimi Vendor Verifier and is ready for production.
Related event: SGLang Day-0 Support for Kimi K3 Boosts Throughput to 423 tok/s(8 posts)→
More from Infra
- Engram's random reads don't suit SSDs; CPU-memory over NVLink could serve all 72 GPUs — bookwormengr · 2026-09-11
- 80% of the DIY LLM inference hype posters have already quit — it's brutally hard systems work — abhijithneil · 2026-09-11
- Hugging Face's Ultra Scale Playbook: a free book on training LLMs on GPU clusters — mdancho84 · 2026-09-11
- Is inference latency becoming the biggest bottleneck for production AI agents? — Euphoric_Sea632 · 2026-09-11
- LLM Serving Metrics Thread: Why TPOT and Uptime Make or Break User Experience — abhijithneil · 2026-09-11
- PlanetScale launches sharded Postgres: 768 servers acting as one, 1PB scale — dhruv2038 · 2026-09-11