Sparse attention is heating up: HISA blends DSA and NSA ideas, may slash indexer costs

teortaxesTex · x · 2026-09-17

Commenting on a new paper, teortaxesTex notes that sparse attention research is moving fast. With DSA (DeepSeek Sparse Attention) now proven effective, 1M-token contexts are "just the appetizer."

The paper's HISA method reportedly improves on prior results effortlessly, combining DSA with the best ideas from NSA (Native Sparse Attention). In the limit it cuts indexer costs by a factor of B, and it's plug-and-play.

Original post →

More from Research

Research channel →