What is Prefill-Decode Disaggregation and Why Are Modern Inference Stacks Moving to It?

scareme_please · reddit · 2026-08-14

Reddit user scaremeplease asks about prefill-decode disaggregation (PD disaggregation). They understand traditional LLM serving as unified model: prompt hits GPU, GPU runs prefill, then same GPU decodes. Providers are splitting these phases across hardware, and they wonder why and how. Requesting breakdown of PD disaggregation and why custom infra is built around it.

Original post →

More from Infra

Infra channel →