vLLM 0.25.0 Released
vllm_project · x · 2026-07-12
vLLM releases v0.25.0, featuring 558 commits and 232 contributors, 64 of whom are new.
Major changes include:
- Model Runner V2 becomes the default execution path for all dense models.
- Legacy PagedAttention implementation removed; V1/MRv2 is now the standard path.
- Transformers backend speed increased to near native vLLM levels.
- Added unified Streaming Parser Engine.
- Supports universal speculative decoding (TLI) across heterogeneous vocabularies, with DSpark and DFlash drafters added.
- Added/updated multiple models like Hy3, Unlimited OCR, LLaVA-OneVision-2, MOSS-Transcribe-Diarize, openai/privacy-filter, GLM-5, DeepSeek-V3.2, and MiniMax-M3.
More from Infra
- Production AI budgets include retries, routing, caching and observability—not just token prices — arx-go · 2026-07-22
- NVIDIA briefs analysts on Vera CPU and doubles down on monolithic agentic design — BenBajarin · 2026-07-22
- NVIDIA unveils Vera Rubin platform with claims of 10x better performance per watt — nvidia · 2026-07-22
- SkyPilot comes out of stealth with a pitch to unify fragmented AI compute — skypilot_org · 2026-07-22
- Why a 1GW Chinese AI data center may be plausible after all — teortaxesTex · 2026-07-22
- AI agent accountability layer adds terminal verification with explicit finality and no signup — Special_Librarian145 · 2026-07-22