Prefill vs Decode: why prompts are compute-bound but generation is bandwidth-bound

ariG23498 · x · 2026-10-09

A new llama.app doc breaks LLM inference into two phases: prefill processes the prompt in parallel and is compute-bound (longer prompts, bigger models and slow kernels hurt), while decode generates tokens autoregressively, repeatedly reading weights and the KV cache from memory, making it memory-bandwidth-bound — batching helps amortize the cost. Recommended by HF engineer Merve as a crisp primer on inference bottlenecks.

Related event: Hugging Face Launches Guide to LLM Inference: Prefill, Decode and KV Cache(3 posts)→

Original post →

More from Infra

Infra channel →