Recent LLMs Adopt Different KV Head Configurations in SWA Layers
Several recent LLM implementations use a higher number of KV heads in sliding window attention (SWA) layers compared to global layers, a detail confirmed across multiple works.
2026-07-16 ~ 2026-07-16 · 2 related posts
- KV Head Count Discrepancy in SWA Layers is Crucial — loiccabannes · 2026-07-16
- More Recent Works Use Different KV Head Configurations — eliebakouch · 2026-07-16