Sliding Window Attention Beats Linear Attention Post-Training

jm_alexia · x · 2026-08-31

Research demonstrates that switching a pretrained softmax transformer to a sliding-window attention mask with attention sinks (SWA) outperforms post-training adaptation to linear attention, requiring no additional training. Commentary suggests that for linear attention to work effectively, the architecture should be pretrained with it, rather than adapted later, as invasive changes to strong pretrained models degrade performance.

Related event: Sliding Window Attention with Sinks Beats Post-Trained Linear Attention Without Training(5 posts)→

Original post →

More from Research

Research channel →