vLLM pushes Kimi K3 serving to 2.2–2.8x throughput, 56–60% lower latency
woosuk_k · x · 2026-09-18
The vLLM team, together with Red Hat, NVIDIA, Huawei, and Inferact, published a deep dive on Kimi K3 serving optimization: versus v0.27.1 on a B300 node (8K/1K workload, TP8), throughput improved 2.2–2.8x, latency dropped 56–60%, and TTFT fell 72–85% across concurrency 1/4/16.
Key optimizations:
- Adaptive speculative-token budgets with 8-token DSpark speculation
- Internal KDA prefix checkpoints and zero-copy mixed KDA batches to fix recurrent-state bottlenecks
- Deferred MXFP4 expert finalization and elimination of small tensor copies
- ReplaySSM, prefill/decode disaggregation, cache offload, and decode context parallelism
The authors note scheduler limits and small tensor copies can matter as much as large GEMMs. Full launch commands and reproduction steps are included; the effort is tracked in issue #50587.
Related event: vLLM Optimizes Kimi K3 Inference with Up to 2.8x Throughput Gain(2 posts)→
More from Infra
- Starlink connects thousands of rural Latin American schools, from 10,000 antennas in Honduras to Bolivia's national rollout — XFreeze · 2026-09-18
- 0.05% sampling to validate cache hits: developer marvels at compute saved across the system — DanielLockyer · 2026-09-18
- TRL Adds Async GRPO with LoRA Weight Sync over HF Buckets, Cutting Training from 3.5h to 53min — _lewtun · 2026-09-18
- Payments firms race to own AI inference: Stripe taps OpenRouter, Ramp enters the chain — xkonjin · 2026-09-18
- OpenDCAI/DataFlow: open-source pipeline toolkit for pre-training data prep — Puzzleheaded_Box2842 · 2026-09-18
- Qdrant wraps 4+ hour Vector Space Stream on vector search — recording now live — qdrant_engine · 2026-09-18