BlockRank: sparse attention makes LLM in-context document ranking faster
dejanseo · x · 2026-10-05
dejan.ai breaks down BlockRank, a UT Austin + Google method tackling the O(n²) attention bottleneck of LLM-based In-Context Ranking using structured sparse attention and contrastive training, boosting ranking speed without accuracy loss. Commenter natzir9 notes the experiments still rerank candidates placed in the prompt, so it fundamentally depends on a retrieval index — and points to a related Google patent storing model predictions inside the index.
Related event: BlockRank Speeds Up LLM Document Ranking with Block-Sparse Attention(2 posts)→
More from Research
- Human-AI collaboration pushes 11-square packing lower bound to 3.875, closing 95.89% of the gap — ctjlewis · 2026-10-05
- Follow-up on the same 11-square packing breakthrough: 300-line verifiable proof via 1,121 weighted dots — ctjlewis · 2026-10-05
- Pivot-SD: 200 questions beat diffusion RL by distilling only high-impact denoising pivots — kaist-ai · 2026-10-05
- SBERT2S1 turns biomedical retrieval encoders into calibrated one-pass typed decision models — Pritam Deka · 2026-10-05
- 315M-param self-play RL bot Pluto crushes GPT-6 Astra 0-1000 in StarCraft: Brood War — i_dg23 · 2026-10-05
- Same system scores 51.4 vs 75.8 depending on judge — are memory benchmarks trustworthy? — Efficient_Joke3384 · 2026-10-05