Inferact’s Kimi-K3-DSpark draft model reuses MLA caches to speed up vLLM serving
vllm_project · x · 2026-07-27
- Inferact’s Kimi-K3-DSpark is an MLA-native draft model for accelerating Kimi K3 inference on vLLM.
- The draft is trained on target hidden states extracted from vLLM itself, so it learns from the exact distribution it will later serve.
- DSpark combines a block-diffusion backbone (5 dense layers drafting 7 tokens in parallel), a rank-256 sequential Markov head, and a confidence head.
- Because the draft mirrors Kimi K3’s MLA attention, draft and target share a single compact KV layout, enabling reuse of the same cache management and PD-disaggregated serving path.
- The card also notes that acceptance is higher on deterministic text like code and lower on high-entropy generation like creative writing, using production-like settings such as tensor-parallel-size=8, temperature=1.0, and topp=0.95.
Related event: vLLM brings day-0 support to Moonshot’s Kimi K3(11 posts)→
More from Infra
- Together AI adds canary rollouts for zero-downtime model upgrades on dedicated inference — togethercompute · 2026-09-23
- Dedicated Hardware for Running AI Agents at Scale Arrives — cyrilzakka · 2026-09-23
- Ternary Bonsai 2 27B: 5.9GB weights retain ~95% of full-precision reasoning — cephaloform · 2026-09-23
- Qwen 27B runs 24hr unattended on one RTX5090, builds full Postgres-SpringBoot-React spreadsheet app — anglepoiselife · 2026-09-23
- OpenRoboto Shift launches: decentralized egocentric video data network for robot brains — markjeffrey · 2026-09-23
- Engineer describes designing digital circuits that recycle most of their energy — MikePFrank · 2026-09-23