Sliding Window Attention Beats Linear Attention, but Hybrid Architectures Reign Supreme

ChengleiSi · x · 2026-09-01

A paper claims that sliding window attention (SWA) with attention sinks outperforms linear attention at no cost. Critics argue this is misleading: linear attention initialized on full-attention weights isn't optimally trained. Furthermore, frontier models favor hybrid architectures (e.g., Kimi's 3:1 ratio), which empirically outperform both pure linear and pure full attention mechanisms.

Related event: Microsoft research: sliding-window attention with sinks beats post-trained linear attention at zero cost(9 posts)→

Original post →

More from Models

Models channel →