vLLM v0.30.0 ships with 762 commits: watermarking, HiSparse, Model Runner V2
vllm_project · x · 2026-09-23
vLLM v0.30.0 is out with 762 commits from 315 contributors (104 first-timers). Highlights:
- Hybrid-attention hot paths: Kimi K3 streamlines KDA/AtRes/MLA; DeepSeek-V4.1-Flash adds MXFP8 KV and async Engram; Qwen3.8-Flash-Next fuses QSA/PLE to cut sparse-GQA overhead
- HiSparse: adds a host tier beneath sparse-MLA decode; under GPU pressure only top-k misses return to a per-request hot buffer
- Model Runner V2: EAGLE3-style drafts for pipeline parallelism, adaptive verification extended to every draft-model speculator via online acceptance estimation
- Dual-key Gumbel-max watermark generation and detection, with per-request opt-out and speculative decoding support
- New models: GLM-5.3-Flash, K2-Horizon, Cohere Compass, Bailing V3 VL
- Fast Start: post-quantized TP-sharded weights resident in a per-GPU daemon, restarts map them via CUDA IPC with --load-format ipccache
Related event: vLLM v0.30.0 Released with 762 Commits and Major Performance Gains(3 posts)→
More from Infra
- Anthropic in early talks to lease up to 1GW from Apollo-backed data center developer — rohanpaul_ai · 2026-09-23
- Anthropic in early talks to lease up to 1GW of compute from Apollo-owned data center operator — rohanpaul_ai · 2026-09-23
- Inference startup: training wins inference workloads, models should improve with use — ypatil125 · 2026-09-23
- Apple M6 ANE memory bandwidth tops 150GB/s, matching GPU/CPU — AIFlow_ML · 2026-09-23
- Report: Anthropic in talks to cement control over more data centers — pstAsiatech · 2026-09-23
- Back-of-envelope math: MiMo's 1.27M RL rollouts could run in ~10 hours — teortaxesTex · 2026-09-23