Sliding-window attention beats quadratic attention in new paper
woadwarrior · reddit · 2026-08-31
Highlighting a new paper from Alexia Jolicoeur-Martineau et al. The research proposes replacing quadratic attention with sliding window attention combined with attention sinks, requiring no post-training. This method could significantly reduce memory usage for local LLM inference on memory-constrained hardware.
More from Infra
- SK hynix reportedly considers Intel Foundry for HBM4E base dies — AccBalanced · 2026-09-01
- Guardian proposes off-grid, self-powered datacenters to cut emissions — nordicinst · 2026-09-01
- Compute Wants to Leave Earth: A Manifesto for Orbital Infrastructure — McDonaghMatthew · 2026-09-01
- Optimizing Qwen 3.8 Flash Next: Improving speeds on 64GB VRAM setup — Jorlen · 2026-09-01
- How much power does the AI buildout actually take? — TheZachMueller · 2026-08-31
- Tested 3 Claude Code plugins to reduce costs: Here is what worked — Marmelab · 2026-08-31