MHA, MQA, GQA and MLA explained by what happens to the K/V cache during decoding

techNmak · x · 2026-09-13

A first-principles walkthrough of attention variants, arguing the real distinction is what happens to the key/value state during decoding.

The companion technical handbook covers single-head attention, the KV-cache formula, decode bandwidth, GQA uptraining, MLA latent compression and projection absorption, decoupled RoPE, worked cache comparisons, FlashAttention and PagedAttention, implementation details, and common mistakes — grounded in the original papers.

Original post →

More from Infra

Infra channel →