Chamath breaks down AI compute: prefill is compute-bound, decode is bandwidth-bound

rohanpaul_ai · x · 2026-09-28

Chamath explains the two key phases of AI compute: Prefill processes the input context and is compute-bound, favoring massively parallel GPUs — which is why Nvidia dominates as context lengths grow. Decode generates tokens one at a time, with each next token requiring a scan of everything generated so far, making it memory-bandwidth bound rather than compute-bound. The takeaway: inference optimization's real battleground is bandwidth, not FLOPs.

Original post →

More from Infra

Infra channel →