Benchmark: Prefill/Decode split fails on commodity hardware

vllm_project · x · 2026-08-29

An open-source benchmark using vLLM + NIXL reveals that splitting LLM prefill and decode phases hurts performance on commodity hardware. Disaggregating across GPUs without fast links results in 5.5× worse TTFT, and CPU prefill is 100× slower. The study concludes this architecture only wins at frontier-lab scale with NVLink/RDMA and high concurrency.

Original post →

More from Infra

Infra channel →