KV Head Count Discrepancy in SWA Layers is Crucial

loiccabannes · x · 2026-07-16

An easily overlooked but potentially crucial implementation detail: some recent works (like Gemma 4, inkling, and the author's own SDM) use a higher number of KV heads in their SWA layers than in their global layers.

The author believes this design choice might directly determine whether the model can outperform GDN/mamba or actually perform worse.

Related event: Recent LLMs Adopt Different KV Head Configurations in SWA Layers(2 posts)→

Original post →