vLLM pushes Kimi K3 serving to 2.2–2.8x throughput, 56–60% lower latency

woosuk_k · x · 2026-09-18

The vLLM team, together with Red Hat, NVIDIA, Huawei, and Inferact, published a deep dive on Kimi K3 serving optimization: versus v0.27.1 on a B300 node (8K/1K workload, TP8), throughput improved 2.2–2.8x, latency dropped 56–60%, and TTFT fell 72–85% across concurrency 1/4/16.

Key optimizations:

The authors note scheduler limits and small tensor copies can matter as much as large GEMMs. Full launch commands and reproduction steps are included; the effort is tracked in issue #50587.

Related event: vLLM Optimizes Kimi K3 Inference with Up to 2.8x Throughput Gain(2 posts)→

Original post →

More from Infra

Infra channel →