Inference engineering deep dive: why prefill/decode split and KV-cache dominate serving

techNmak · x · 2026-10-07

A long-form argument that inference engineering is an underrated but critical AI skill set, breaking down the prefill/decode split and the techniques built around it.

Core split

Techniques

Scheduling

Memory: KV cache size depends on layers, cached tokens, KV heads, head dim and bytes per value — which is why MQA/GQA matter so much for serving.

Original post →

More from Infra

Infra channel →