vLLM v0.31.0 ships 717 commits: vllm preload keeps quantized weights in GPU memory across restarts

lmoroney · x · 2026-10-07

vLLM v0.31.0 landed with 717 commits from 307 contributors, targeting the wait of reloading large quantized weights after restarts:

Advice: audit launch scripts and request code for affected flags before upgrading production servers.

Original post →

More from Infra

Infra channel →