6 serving-side techniques that make LLM inference faster - from prefix caching to PD disaggregation

techNmak · x · 2026-09-23

A detailed thread explains how much LLM serving performance comes from around the model, not the model itself, breaking down six techniques that each attack a different source of waste:

Original post →

More from Infra

Infra channel →