vLLM v0.29.0 cuts Blackwell E2E latency 33.6%, with 6.6-7.6x kernel speedups for Kimi-K3
vllm_project · x · 2026-09-11
vLLM's performance and hardware highlights for v0.29.0:
- Kimi-K3: Mamba metadata prep fused into a single Triton launch for 6.67.6x kernel-level speedup
- Batch invariance: tuned per architecture, 3x faster decode kernels on RTX 4090D / H20
- Blackwell autotuning: 33.6% E2E latency reduction
- NVIDIA: DeepSeek V3.2 and GLM-5.2 DSA reach the CUDA implementation on every GPU
- AMD ROCm: W4A4 now defaults to preshuffled asm GEMM; the ROCr/CLR base image fixes graph replay segfaults, up to 20% TPOT improvement at concurrency 1
- CPU: a new AMX MLA backend lets DeepSeek V2/V3/R1 run end to end
- Intel XPU: INC int4 W4A8 linear backend and AutoRound MXFP8 MoE support
Related event: vLLM v0.29.0 Cuts Blackwell Latency 33.6%, Model Runner V2 Default(4 posts)→
More from Infra
- SF Compute signs $245M in take-or-pay contracts for NVIDIA Blackwell B300 capacity — mattshumer_ · 2026-09-11
- SpaceX CFO: vertical integration is core, Starship paves way for orbital compute — elonmusk · 2026-09-11
- Bezos: Power Supply Chain Bottleneck Forces AI Labs to Slow Development Pace — beffjezos · 2026-09-11
- One cheeseburger emits as much CO2 as 63,000 Gemini text prompts, math shows — recallingmemories · 2026-09-11
- Google signs deal to buy half the electricity of a nuclear power plant — lukaspetersson · 2026-09-11
- spcx reportedly signed another mega compute deal a week ago, $13B ARR per CFO — rwang07 · 2026-09-11