Sliding-window attention beats post-trained linear attention

jm_alexia · x · 2026-08-31

A new paper shows that Sliding Window Attention (SWA) with sinks matches or outperforms post-trained Linear Attention models at no cost. SWA achieves 2-10x higher performance on long-context reasoning tasks. It requires no post-training, is fast, and low-memory. The authors recommend switching to SWA over retrofitting linear attention to reduce inference memory costs.

Related event: Study: Sliding Window Attention Beats Linear Attention(3 posts)→

Original post →

More from Research

Research channel →