Inferact’s Kimi-K3-DSpark draft model reuses MLA caches to speed up vLLM serving

vllm_project · x · 2026-07-27

Original post →

More from Infra

Infra channel →