DeepSeek-V4.1-Flash on vLLM: 5.3x agentic throughput three weeks after day 0

vllm_project · x · 2026-10-08

The vLLM team details three weeks of optimization for DeepSeek-V4.1-Flash: 1.9x faster at low concurrency and 5.3x throughput at 150 TPS/user on SemiAnalysis AgentX. Key wins: SWA bounded replay with CUDA graphs (30% TTFT cut), DeepSeek's new kernels (MegaAttention with NVFP4 KV 45% smaller, Mega-mHC, Mega-Gate, DeepSelect), plus aggressive kernel fusion including sparse MQA logits 14-23x faster per layer at 512K context.

Original post →

More from Infra

Infra channel →