vLLM v0.30.0 Released with 762 Commits and Major Performance Gains

vLLM has released v0.30.0 with 762 commits from 315 contributors, adding DeepSeek-V4.1-Flash support and quantifiable gains including a 3.23x speedup for Gemma 4 multimodal inference and improved RL sampling throughput.

2026-09-23 ~ 2026-09-23 · 3 related posts