Inference engineering deep dive: why prefill/decode split and KV-cache dominate serving
techNmak · x · 2026-10-07
A long-form argument that inference engineering is an underrated but critical AI skill set, breaking down the prefill/decode split and the techniques built around it.
Core split
- Prefill: parallel over the prompt, high arithmetic intensity, compute-bound
- Decode: one token per step, repeatedly reads weights and KV cache, memory-bandwidth-bound at practical batch sizes
Techniques
- FlashAttention: fewer HBM↔SRAM reads/writes for exact attention
- PagedAttention: block-based KV-cache memory management, cutting fragmentation
- MQA/GQA: fewer KV heads mean smaller cache per token and less data moved
Scheduling
- Continuous batching: requests join/leave the batch dynamically
- Chunked prefill: splits large prefills to schedule alongside decode, controlling interference
- Prefix caching: reuses KV for shared prefixes — saves prefill compute but doesn't speed up decoding
Memory: KV cache size depends on layers, cached tokens, KV heads, head dim and bytes per value — which is why MQA/GQA matter so much for serving.
More from Infra
- Dev ships Vulkan-only local generative art app with no telemetry or cloud — ogimaru · 2026-10-07
- AMD ships ROCm 10.1, targeting storage-to-GPU data movement as the new training bottleneck — AccBalanced · 2026-10-07
- Crusoe's Path: From Stranded-Gas Bitcoin Mining to Prefab Gigawatt-Scale GPU Datacenters — AccBalanced · 2026-10-07
- Intel to keep working with Elon Musk on Terafab push into cutting-edge chips — SumitGup · 2026-10-07
- Musk bets every 5GW of added US power equals roughly 1% GDP growth — XFreeze · 2026-10-07
- Burn 0.22 released: biggest Rust DL framework update yet, 6-15x faster rebuilds — JosephJacks_ · 2026-10-07