S2-Attention: hardware-aware Triton kernels make sparse attention actually fast
burkov · x · 2026-09-11
Dense attention's quadratic cost bottlenecks LLM training and inference. Sparse attention cuts theoretical FLOPs but rarely delivers wall-clock speedups due to missing hardware-level memory optimizations, and often loses accuracy on long-context tasks; inference-time token eviction also fragments memory in serving frameworks like PagedAttention.
The paper introduces Sparsely-Sharded Attention (S2-Attention), a hardware-aware Triton kernel library achieving real acceleration without accuracy loss. Its Merge-Q technique dynamically merges query blocks sharing key-value data into single computation tiles, cutting redundant loads and turning sparsity into actual speedups.
More from Infra
- Colocated async RL gains steam as observers speculate k3 uses it too — stochasticchasm · 2026-09-11
- Peter Diamandis: The AI race is becoming the biggest construction project of our generation — PeterDiamandis · 2026-09-11
- k3 Report Section Confirms Millions of Concurrent Sandboxes in Its RL Training Run — stochasticchasm · 2026-09-11
- Pentagon in talks to lend roughly $5 billion to AI cloud startup Fluidstack — vitaliychiley · 2026-09-11
- Eric Schmidt: AI may hit a money wall before a power wall — $1T capital needed — rohanpaul_ai · 2026-09-11
- SpaceX signs another AI compute deal: $1.11B per month, on track for $100B ARR — NinaDSchick · 2026-09-11