Long-Short Sliding Window Attention trims 3 ms per step in long-context training

gordic_aleksa · x · 2026-07-27

The author explains why a particular long-context structure can hurt performance at scale, especially as sequence length grows.

Core idea

Training trick

Why it matters

Original post →

More from Infra

Infra channel →