Long-Short Sliding Window Attention trims 3 ms per step in long-context training
gordic_aleksa · x · 2026-07-27
The author explains why a particular long-context structure can hurt performance at scale, especially as sequence length grows.
Core idea
- The first technique is Long-Short Sliding Window Attention.
- It is inspired by Gemma 2’s global-local style and other hybrid architectures, but differs in two ways:
- It uses sliding-window attention instead of mixing full attention with linear attention layers.
- Some layers use a long SWA context while others use a short SWA context.
Training trick
- The context length of the long SWA layers is warmed up at double the rate of the short SWA layers.
- This reduced step time by about 3 ms/step.
- After adding 15 extra training steps to offset the slightly worse validation loss, the net gain was still about 2–3 seconds overall.
Why it matters
- The post argues that this kind of structure can improve speed, but may not scale cleanly to very long contexts.
- The attached charts show both the layer pattern and how attention entropy changes as the context length warms up.
More from Infra
- Triton backend pushes Falcon3-10B to 97.5 tok/s on an RTX 5070 — OCV_Researcher · 2026-07-27
- Sparrow switches its Standard mode to Ministral 3 14B for local document extraction — andrejusb · 2026-07-27
- Nvidia supplier Wistron opens $700 million Texas plant for GB300 and Vera Rubin systems — Beth_Kindig · 2026-07-27
- BeeLlama.cpp v0.4.1 adds KV-cache precision tails and new quantization modes — Anbeeld · 2026-07-27
- Running 100 million tokens through GLM 5.2 NVFP4 locally costs about $1 — _akhaliq · 2026-07-27
- Cloudflare’s AI-training block can also stop Googlebot after September 15 — daluoseo · 2026-07-27