Learnable soft top-K sparsification cuts LLM sparse retrieval latency, accepted at SIGIR 2026

_reachsumit · x · 2026-10-05

An arXiv paper (2610.02572, accepted at SIGIR 2026) tackles a key inefficiency in neural sparse retrieval: LLM term expansion produces overly long sparse vectors due to large vocabularies. The proposed scheme combines learnable soft top-K, per-term thresholding, and FLOPs regularization. On MS MARCO and BEIR with the Lion-SP model, it significantly shortens average query and document lengths, cutting retrieval latency and storage while keeping relevance highly competitive.

Original post →

More from Research

Research channel →