DeepSeek-V4.1-Flash on vLLM: 5.3x agentic throughput three weeks after day 0
vllm_project · x · 2026-10-08
The vLLM team details three weeks of optimization for DeepSeek-V4.1-Flash: 1.9x faster at low concurrency and 5.3x throughput at 150 TPS/user on SemiAnalysis AgentX. Key wins: SWA bounded replay with CUDA graphs (30% TTFT cut), DeepSeek's new kernels (MegaAttention with NVFP4 KV 45% smaller, Mega-mHC, Mega-Gate, DeepSelect), plus aggressive kernel fusion including sparse MQA logits 14-23x faster per layer at 512K context.
More from Infra
- Nebius up 160% vs CoreWeave's 10%: the AI infrastructure stock divergence explained — Beth_Kindig · 2026-10-08
- Microsoft's $5,999 Surface RTX Spark Dev Box preorders open, ships November — tomwarren · 2026-10-08
- Together scales open-source inference with IBM and NVIDIA on B300 cluster — togethercompute · 2026-10-08
- How a Self-Appending Summarizer Quadrupled Token Costs in a 1.4M-Conversation Agent — Good_Education4713 · 2026-10-08
- National Compute gifts $100M in compute credits to White House Genesis Mission — typewriters · 2026-10-08
- Ask: Is Qwen 3.8 Flash-Next Worth It Over 3.6 35B-A3B on 64GB RAM + 16GB VRAM? — msalsas · 2026-10-08