vLLM Ships Full Support for DeepSeek-V4.1-Flash's New Architecture
vllm_project · x · 2026-09-10
The vLLM project announced full support for DeepSeek-V4.1-Flash's new architecture, including:
- Multiple parallelisms: TP/SP/DP/EP
- DSpark speculative decoding with adaptive verification
- Agentic-first features: prefix caching, KV cache offloading, and PD disaggregation
- Engram CPU offloading plus DP-aware sharding
- Compute- and KV-cache-efficient SWA bounded replay
More kernel integrations from DeepGeMM and FlashMLA are landing soon.
More from Infra
- Google Cloud user hit with an $82k bill within 5 hours — Patient_Election2179 · 2026-09-10
- TRL ships 1M-token long-context training guide, trains Qwen3-8B on one 8-GPU node — QGallouedec · 2026-09-10
- Dual RTX Pro 6000 + Threadripper 9955W local LLM build — sanity check requested — No_Run8812 · 2026-09-10
- Screenshot surfaces rare admission of 72-hour KV cache limits in V4-era architecture — zephyr_z9 · 2026-09-10
- DeepSeek cut per-token KV cache size by 54x in nine months — zephyr_z9 · 2026-09-10
- No one matches its inference economics; commentator suggests 10-20% cut from infra providers — zephyr_z9 · 2026-09-10