vLLM cuts TTFT nearly 70% at ~100K throughput with DeepSeek NVFP4 kernels and fused ops
vllm_project · x · 2026-10-08
vLLM announced a stack of inference optimizations delivering nearly 70% lower TTFT at 100K throughput.
- Kernels: integration of DeepSeek's MegaAttention (NVFP4 KV cache, 45% smaller), Mega-mHC, Mega-Gate and DeepSelect.
- vLLM fusions: a CuTe-DSL fused WO-A op (up to 6–7% lower ITL), mHC coefficients on a side stream, and sparse MQA logits (14–23× faster per layer at 512K context).
Built by @inferact and the vLLM community, with models/kernels from DeepSeek, collaboration from NVIDIA, and AgentX support from SemiAnalysis.
More from Infra
- xAI Bets on Leasing Data Centers Over Owning a Frontier Lab, Says Investor — PaulYacoubian · 2026-10-08
- Tech Drives 3/4 of S&P 500 Earnings Growth as AI Trade Disperses — demian_ai · 2026-10-08
- OpenAI's 26M-Line AI Proof Crashed Linux's Default mmap Limit — ceciletamura · 2026-10-08
- Mistral researcher: 15T tokens underestimate pretraining, closer to 40T+ — eliebakouch · 2026-10-08
- Higgsfield at DevDay 2026: Agents that provision their own compute — OpenAI · 2026-10-08
- Nebius up 160% vs CoreWeave's 10%: the AI infrastructure stock divergence explained — Beth_Kindig · 2026-10-08