KV Head Count Discrepancy in SWA Layers is Crucial
loiccabannes · x · 2026-07-16
An easily overlooked but potentially crucial implementation detail: some recent works (like Gemma 4, inkling, and the author's own SDM) use a higher number of KV heads in their SWA layers than in their global layers.
The author believes this design choice might directly determine whether the model can outperform GDN/mamba or actually perform worse.
Related event: Recent LLMs Adopt Different KV Head Configurations in SWA Layers(2 posts)→