LLM decoding explained: prefill reads in parallel, decode writes token by token

abhijithneil · x · 2026-09-04

Parts 5-6 of @abhijithneil's LLM inference primer explain the two core stages:

Foundational concepts for understanding inference latency and throughput metrics.

Related event: Viral Thread Explains LLM Inference Metrics: TTFT, TPS and TPOT(4 posts)→

Original post →

More from Infra

Infra channel →