vLLM v0.31.0 ships 717 commits: preload daemon keeps quantized weights resident across restarts
AccBalanced · x · 2026-10-07
vLLM v0.31.0 landed with 717 commits from 307 contributors, targeting the pain of reloading large quantized weights after every server restart.
- vllm preload: a new command that runs a daemon keeping post-quantized weights resident in GPU memory across engine restarts, with a /health endpoint and readiness wait
- Experimental CRIU snapshots: restore a fully initialized engine running on one GPU, skipping cold start
- Breaking changes: per-request mmprocessorkwargs and mediaiokwargs are now rejected unless --trust-request-mm-kwargs is set; tokenizermode="slow" is removed — worth reviewing the full list before upgrading production servers
Related event: vLLM v0.31.0 Ships CRIU Snapshots and Preload to Kill Cold Starts(3 posts)→
More from coding & agent
- Managed Agents on Interactions API: one call spins up Antigravity agents in a remote sandbox — clmt · 2026-10-08
- Legacy billing rewrite estimated at 7-8 months shipped in six, 122 PRs in 90 days — alex_verem · 2026-10-08
- LangChain engineer built an ACP coding agent that replaced Claude Code for 9 months — Hacubu · 2026-10-08
- 30 Real Business Workflow Tests: Keep Agent Evaluation Simple — VibeMarketer_ · 2026-10-08
- Hybrid agent pattern: cloud Gemini plans, local Gemma swarm runs 97% of tokens offline — clmt · 2026-10-08
- Long-running agents suffer 'constraint amplification': a subtle form of context rot — generativist · 2026-10-08