Sliding-Window Attention Beats Linear Attention in Long-Context Tasks

iScienceLuvr · x · 2026-08-31

Research shows that Sliding Window Attention (SWA) with sinks performs as well as or better than post-trained Linear Attention models across multiple LLMs and downstream tasks. In long-context reasoning tasks like Needle-in-a-Haystack and BABILong, SWA achieves 2 to 10 times higher performance. SWA requires no post-training, is extremely fast, and needs low memory, making it a cost-effective solution for reducing inference costs.

Related event: Sliding Window Attention with Sinks Beats Post-Trained Linear Attention Without Training(5 posts)→

Original post →

More from Research

Research channel →