vLLM v0.29.0 Cuts Blackwell Latency 33.6%, Model Runner V2 Default
vLLM v0.29.0 ships with 594 commits from 277 contributors, cutting Blackwell end-to-end latency by 33.6% and making Model Runner V2 the default, alongside KV offloading, queue controls, and removal of ten deprecated architectures.
2026-09-11 ~ 2026-09-11 · 4 related posts
- vLLM v0.29.0 ships Model Runner V2 as default with 594 commits from 277 contributors — vllm_project · 2026-09-11
- vLLM v0.29.0 cuts Blackwell E2E latency 33.6%, with 6.6-7.6x kernel speedups for Kimi-K3 — vllm_project · 2026-09-11
- vLLM upgrade guide: KV offloading, queue admission control, 33.6% Blackwell latency cut — vllm_project · 2026-09-11
- vLLM v0.29.0 upgrade guide: new defaults, ten architectures removed, DoS fix — vllm_project · 2026-09-11