BlockRank: sparse attention makes LLM in-context document ranking faster

dejanseo · x · 2026-10-05

dejan.ai breaks down BlockRank, a UT Austin + Google method tackling the O(n²) attention bottleneck of LLM-based In-Context Ranking using structured sparse attention and contrastive training, boosting ranking speed without accuracy loss. Commenter natzir9 notes the experiments still rerank candidates placed in the prompt, so it fundamentally depends on a retrieval index — and points to a related Google patent storing model predictions inside the index.

Related event: BlockRank Speeds Up LLM Document Ranking with Block-Sparse Attention(2 posts)→

Original post →

More from Research

Research channel →