vLLM v0.29.0 ships Model Runner V2 as default with 594 commits from 277 contributors
vllm_project · x · 2026-09-11
vLLM released v0.29.0 with 594 commits from 277 contributors (91 first-timers). Key changes:
- Model Runner V2 is now the default runner for all models
- Another optimization round for Kimi-K3, with the latent tail lowered to Mamba metadata
- DeepSeek-V4 shared experts fused into MegaMoE
- Mamba prefix caching keeps internal prefill checkpoints: 925% better TTFT
- Speculative decoding reports per-request acceptance stats over the API
- RL weight sync adds the shardedrdt P2P backend so each worker pulls only its own TP/EP slice
- New models: Hy4-preview, Qwen3.8-Flash-Next, GraniteSWA, NemotronH Omni Reasoning V3
Related event: vLLM v0.29.0 Cuts Blackwell Latency 33.6%, Model Runner V2 Default(4 posts)→
More from Infra
- REVA Mines LLM Attention into Reusable Evidence Views, Cutting RAG Compression Overhead up to 15.6x — _reachsumit · 2026-09-11
- Edge0-35B-A3B preview MoE model for edge inference trends on Hugging Face — Edge0 · 2026-09-11
- Reflect Orbital wants to sell sunlight via volleyball-court mirrors on satellites — kyliebytes · 2026-09-11
- Data centers are for startups, not frontier labs: more compute is the anti-monopoly move — arthurcolle · 2026-09-11
- DeepSeek cut KV cache per token 54x in 9 months, called the third frontier lab — max_paperclips · 2026-09-11
- Anatomy of Jensen Huang's 'AGI Is Here' Tweet: Zero-Cost Signaling and Abilene's 357-Job Data Center Deal — 创业邦 · 2026-09-11