vLLM v0.31.0 ships 717 commits: vllm preload keeps quantized weights in GPU memory across restarts
lmoroney · x · 2026-10-07
vLLM v0.31.0 landed with 717 commits from 307 contributors, targeting the wait of reloading large quantized weights after restarts:
- vllm preload: a daemon keeps post-quantized weights resident in GPU memory across engine restarts, with a /health endpoint and readiness wait.
- Experimental snapshots via CRIU can restore a fully initialized engine on one GPU.
- Breaking changes: quantization="fp8" for online quantization is replaced by fp8pertensor; tokenizermode="slow" removed; per-request mmprocessorkwargs and mediaiokwargs are rejected unless --trust-request-mm-kwargs is set.
Advice: audit launch scripts and request code for affected flags before upgrading production servers.
More from Infra
- Serving & Monitoring Local LLMs Across Three Mixed-GPU Machines Without Duct Tape — ziyaulhuk12 · 2026-10-07
- Disaggregated inference is the future, says e/acc's Beff Jezos after panel with General Compute — beffjezos · 2026-10-07
- What Does 'Owning Your Own AI' Technically Mean? Locally Running Open Weights Debated — Imaginary_Choice_430 · 2026-10-07
- NVIDIA's LoGRA cuts LLM RL training memory by up to 45.7%, enables 27B RL on a single 8-GPU node — nvidia · 2026-10-07
- Ex-NVIDIA engineer tells the story of testing a chip with a broken memory controller — blelbach · 2026-10-07
- Follow the money risk: memory suppliers take deposits, AI chip makers finance their customers — tengyanAI · 2026-10-07