Kimi K3 serving stack reaches 423 tok/s after DSpark draft-model tuning
ying11231 · x · 2026-07-28
Radixark says they trained a DSpark speculator draft model for Kimi K3 with SpecForge, raising batch-1 decode throughput from about 113 tok/s to 423 tok/s.
Reported gains
- +68% throughput at bs=256 on chat vs. verify-all, lossless
- Accept length of 5.67 on GSM8K and 5.36 on HumanEval
The quoted SGLang update says K3’s fast serving comes from native implementation and optimization of the new architecture, including fused KDA decode kernels, DP attention, DSpark, PD disagg, and KDA-aware prefix caching. It also says the stack has passed Kimi Vendor Verifier and is ready for production.
Related event: Kimi K3 Hits 423 tok/s with SGLang Integration(3 posts)→
More from Infra
- SkyPilot says serving Kimi K3 needs multi-node inference and a full stack — skypilot_org · 2026-07-28
- Fireworks says Kimi K3 matches Opus 5 quality at 2x–4.6x lower task cost — lqiao · 2026-07-28
- Muon orthogonalization goes peer-to-peer to cut all-gather overhead — stochasticchasm · 2026-07-28
- SGLang Day-0 Support for Kimi K3 Hits 423 tok/s on GSM8K — ying11231 · 2026-07-28
- MoE training paper proposes workload-aware expert-GEMM scheduling — stochasticchasm · 2026-07-28
- MoonEP keeps MoE training perfectly balanced with dynamic redundant experts — stochasticchasm · 2026-07-28