vLLM Releases Version 0.25.0
vllm_project · x · 2026-07-12
vLLM officially releases v0.25.0, a massive update comprising 558 commits and 232 contributors.
Key highlights:
- Model Runner V2 becomes the default execution path for all dense models.
- Legacy PagedAttention is retired.
- Transformers backend speed is now close to native vLLM.
- Introduced a unified Streaming Parser Engine.
- Added universal speculative decoding (TLI) across heterogeneous vocabularies, alongside new DSpark and DFlash drafters.
- New model support including Hy3 and Unlimited OCR.
This is the main release announcement for the vLLM upgrade, which subsequent detailed serving and hardware notes are based upon.
More from Infra
- Production AI budgets include retries, routing, caching and observability—not just token prices — arx-go · 2026-07-22
- NVIDIA briefs analysts on Vera CPU and doubles down on monolithic agentic design — BenBajarin · 2026-07-22
- NVIDIA unveils Vera Rubin platform with claims of 10x better performance per watt — nvidia · 2026-07-22
- SkyPilot comes out of stealth with a pitch to unify fragmented AI compute — skypilot_org · 2026-07-22
- Why a 1GW Chinese AI data center may be plausible after all — teortaxesTex · 2026-07-22
- AI agent accountability layer adds terminal verification with explicit finality and no signup — Special_Librarian145 · 2026-07-22