Microsoft Paper: Sliding Window Attention Beats Linear Attention for Inference Memory
rohanpaul_ai · x · 2026-09-02
A new Microsoft paper finds that for reducing inference memory, training-free Sliding Window Attention (SWA) outperforms most retrofitted linear-attention methods.
Methodology: Keep only a small recent window plus the first 4 “sink” tokens that models rely on.
Key Findings:
- With a 64-token window, this setup achieved the best average downstream score in 9 of 11 model comparisons and recovered 99.0% of the full-attention baseline.
- The performance gap widened on long-context reasoning: at 4K context, SWA reached 17.2%–23.0% on Needle-in-a-Haystack tasks (vs. 5.8% for LoLCATs) and 15% on BABILong (vs. 3%).
- In speed and memory tests, the 64-token SWA setup was the fastest and most memory-efficient.
Conclusion: While full attention still wins on very long contexts, SWA with attention sinks is recommended as the first try for fixed, low-memory scenarios without retraining.
More from Research
- Analysis of ExploitGym: OpenAI Model Used Specific Vulnerabilities for Hacking — BlackHC · 2026-09-02
- Lapis (ECCV 2026) Open Sourced: Pixel-Space Diffusion Depth Estimation with Linear Attention — kwangmoo_yi · 2026-09-02
- Lapis: Efficient and High-Quality Depth Estimation via Pixel-Space Diffusion with Linear Attention — kwangmoo_yi · 2026-09-02
- MiniMax Releases H3-World: An Interactive World Model — CryptoBeth96 · 2026-09-02
- Open Source DCLAP Model and SAE Analysis Tool to Fix Long-Tail Terms in Music Search — Old_Rock_9457 · 2026-09-02
- Dev notes record outcomes, but the reasoning dies in the transcript — Sea-Perception1619 · 2026-09-02