vLLM v0.29.0 Cuts Blackwell Latency 33.6%, Model Runner V2 Default

vLLM v0.29.0 ships with 594 commits from 277 contributors, cutting Blackwell end-to-end latency by 33.6% and making Model Runner V2 the default, alongside KV offloading, queue controls, and removal of ten deprecated architectures.

2026-09-11 ~ 2026-09-11 · 4 related posts