vLLM upgrade guide: KV offloading, queue admission control, 33.6% Blackwell latency cut
vllm_project · x · 2026-09-11
The vLLM project published a pre-upgrade rundown of serving, frontend, and performance changes:
- Serving/frontend: Mooncake Store can now offload decode KV; new queue admission control flags (--max-num-queued-reqs, --max-num-queued-tokens); a /v1/messages/render endpoint for the Anthropic Messages API; Rust frontend adds audio/video inputs over gRPC; bounded cachesalt length fixes a scheduler CPU-exhaustion DoS.
- Performance: Kimi-K3 Mamba metadata prep fused into a single Triton launch (6.6–7.6x at kernel level); per-architecture batch invariance yields 3x decode kernels on RTX 4090D / H20; Blackwell autotuning cuts end-to-end latency 33.6%.
- Hardware: NVIDIA reports DeepSeek V3.2 and GLM-5.2 DSA reaching CUDA implementations on every GPU; AMD ROCm defaults W4A4 to the preshuffled asm GEMM.
Related event: vLLM v0.29.0 Cuts Blackwell Latency 33.6%, Model Runner V2 Default(4 posts)→
More from Infra
- SpaceX lands $1.1B/month AI compute hosting deal starting December 2026 — MickeySteamboat · 2026-09-11
- SF Compute signs $245M in take-or-pay contracts for NVIDIA Blackwell B300 capacity — mattshumer_ · 2026-09-11
- SpaceX CFO: vertical integration is core, Starship paves way for orbital compute — elonmusk · 2026-09-11
- Bezos: Power Supply Chain Bottleneck Forces AI Labs to Slow Development Pace — beffjezos · 2026-09-11
- vLLM v0.29.0 cuts Blackwell E2E latency 33.6%, with 6.6-7.6x kernel speedups for Kimi-K3 — vllm_project · 2026-09-11
- One cheeseburger emits as much CO2 as 63,000 Gemini text prompts, math shows — recallingmemories · 2026-09-11