Kimi K3 serving stack reaches 423 tok/s after DSpark draft-model tuning

ying11231 · x · 2026-07-28

Radixark says they trained a DSpark speculator draft model for Kimi K3 with SpecForge, raising batch-1 decode throughput from about 113 tok/s to 423 tok/s.

Reported gains

The quoted SGLang update says K3’s fast serving comes from native implementation and optimization of the new architecture, including fused KDA decode kernels, DP attention, DSpark, PD disagg, and KDA-aware prefix caching. It also says the stack has passed Kimi Vendor Verifier and is ready for production.

Related event: Kimi K3 Hits 423 tok/s with SGLang Integration(3 posts)→

Original post →

More from Infra

Infra channel →