vLLM v0.30.0 Released with 762 Commits and Major Performance Gains
vLLM has released v0.30.0 with 762 commits from 315 contributors, adding DeepSeek-V4.1-Flash support and quantifiable gains including a 3.23x speedup for Gemma 4 multimodal inference and improved RL sampling throughput.
2026-09-23 ~ 2026-09-23 · 3 related posts
- vLLM v0.30.0 ships with 762 commits: watermarking, HiSparse, Model Runner V2 — vllm_project · 2026-09-23
- vLLM v0.30 perf: 3.23x Gemma 4 multimodal speedup, ROCm TPOT -26.4% — vllm_project · 2026-09-23
- vLLM v0.30 ships 762 commits: DeepSeek-V4.1-Flash support and 3.23x Gemma 4 multimodal speedup — vllm_project · 2026-09-23