NVIDIA details EPD disaggregation: up to 5x faster TTFT and 7x faster responses for multimodal serving

NVIDIAAI · x · 2026-09-11

An NVIDIA technical blog explains Encode-Prefill-Decode (EPD) disaggregation in Dynamo, which separates vision encoding from LLM prefill and decode stages. Key results: up to 5x faster time-to-first-token and 7x faster end-to-end latency for image-heavy prompts, short-to-medium outputs, and quantized MoE models; 42.2%/30.8% mean TTFT reduction for text/image requests in mixed traffic by eliminating head-of-line blocking; NVFP4-quantized LLMs with BF16 vision encoders raise colocated EPD goodput from 1.78x to 2.64x. Gains depend on media load, output length, model size/precision, and traffic mix. Includes a GitHub guide plus tips on parallel media decoding, embedding caching, and multimodal KV routing.

Original post →

More from Infra

Infra channel →