vLLM boosts Kimi K3 serving throughput 2.2-2.8x with scheduler, KDA and MoE kernel optimizations

vllm_project · x · 2026-09-17

The vLLM team published a deep dive on optimizing Kimi K3 serving: on a B300 node with an 8K/1K workload at TP8, throughput improved 2.2–2.8x vs v0.27.1, latency dropped 56–60% and TTFT fell 72–85% across concurrency 1/4/16.

Key changes span the full stack:

The post notes scheduler limits and small tensor copies mattered as much as large GEMMs, and ships reproducible server configs and benchmarks. Work is tracked in issue #50587.

Original post →

More from Infra

Infra channel →