vLLM cuts TTFT nearly 70% at ~100K throughput with DeepSeek NVFP4 kernels and fused ops

vllm_project · x · 2026-10-08

vLLM announced a stack of inference optimizations delivering nearly 70% lower TTFT at 100K throughput.

Built by @inferact and the vLLM community, with models/kernels from DeepSeek, collaboration from NVIDIA, and AgentX support from SemiAnalysis.

Original post →

More from Infra

Infra channel →