vLLM v0.30 ships 762 commits: DeepSeek-V4.1-Flash support and 3.23x Gemma 4 multimodal speedup

vllm_project · x · 2026-09-23

vLLM released v0.30.0 with 762 commits from 315 contributors (104 first-timers).

New models

Performance

Fast Start: a persistent per-GPU weight-cache daemon maps post-quantized, TP-sharded weights via CUDA IPC (--load-format ipccache), now covering FP4 and multi-node TP.

Upgrade notes: scale-out endpoints are now opt-in via --enable-scale-out, GPTQ gidx removed, Mamba cache deprecated, YaRN may reduce maxmodellen, default audio resampler is now torchaudio. Also adds Gumbel-max watermarked generation and detection.

Related event: vLLM v0.30.0 Released with 762 Commits and Major Performance Gains(3 posts)→

Original post →

More from Infra

Infra channel →