vLLM v0.28.0 released: 584 commits from 270 contributors
vllm_project · x · 2026-08-27
vLLM v0.28.0 is out with 584 commits from 270 contributors (76 new).
Highlights:
- Stack-wide optimization push for Kimi-K3
- DeepSeek-V4 sparse MLA now works end to end for plain decode, MTP and DSpark
- Speculative decoding adds DFlash2 and DSpark confidence-scheduled verification
- Model Runner V2 picks up E/P/D disaggregation and weight offloading
- Tiered KV offloading gains a disk tier and out-of-tree secondary tiers
- New models: Muse Glimmer, Ling 3.0 Flash, Dots3 NOTE, Interns2mobius
Related event: vLLM v0.28.0 Ships with 584 Commits from 270 Contributors(3 posts)→
More from Infra
- Two vLLM recipes for Blackwell: NVFP4 KV cache buys 262K context and more streams — SeanHighness · 2026-08-27
- Dev forks Nvidia drivers to enable PCIe P2P on GeForce for SlimServe — QuixiAI · 2026-08-27
- View: Single dev with 200 B300s could beat Alibaba's post-training team — kalomaze · 2026-08-27
- GLM 5.3 Flash Benchmark: Hits 881 tok/s on Dual DGX — teortaxesTex · 2026-08-27
- Rakyll advises: If you have any CPU nodes, hold on to them — rakyll · 2026-08-27
- OpenRouter Serves 200B Total Tokens, Adds Qwen 3.5 35B — gajesh · 2026-08-27