Sparse attention is heating up: HISA blends DSA and NSA ideas, may slash indexer costs
teortaxesTex · x · 2026-09-17
Commenting on a new paper, teortaxesTex notes that sparse attention research is moving fast. With DSA (DeepSeek Sparse Attention) now proven effective, 1M-token contexts are "just the appetizer."
The paper's HISA method reportedly improves on prior results effortlessly, combining DSA with the best ideas from NSA (Native Sparse Attention). In the limit it cuts indexer costs by a factor of B, and it's plug-and-play.
More from Research
- A proposal for the future of scientific communication in the age of agents — tensorqt · 2026-09-17
- GPT-Policy: In-Context Robot Learning with VLM Agents, No Gradient Updates — Dongzhou Cheng · 2026-09-17
- Fathom Speeds Million-Token KV Scans 1.67x with Per-Query Read Depth — Vivek Kalyanarangan · 2026-09-17
- nnU-Net Generalization Test on BraTS-GoAT: Dice Drops from 0.906 to 0.831 Across Populations — Tristan Kirscher · 2026-09-17
- Edge0 Streams a 35B MoE from SSD at 20 tok/s on a Single 24GB GPU, Open Source — Edge0 · 2026-09-17
- AI slop's real cost: literature reviews shrinking to under a year of scientific memory — RexDouglass · 2026-09-17