Sparse Attention with Indexer: Distilling Dense Scores

nrehiew_ · x · 2026-08-27

Describes Sparse Attention with an Indexer used at the block level in a compressed latent space. During training, the indexer predicts full attention scores: 1. Distill dense attention scores into the indexer. 2. Apply KL loss between the indexer and the full attention teacher.

Original post →

More from Research

Research channel →