Learnable soft top-K sparsification cuts LLM sparse retrieval latency, accepted at SIGIR 2026
_reachsumit · x · 2026-10-05
An arXiv paper (2610.02572, accepted at SIGIR 2026) tackles a key inefficiency in neural sparse retrieval: LLM term expansion produces overly long sparse vectors due to large vocabularies. The proposed scheme combines learnable soft top-K, per-term thresholding, and FLOPs regularization. On MS MARCO and BEIR with the Lion-SP model, it significantly shortens average query and document lengths, cutting retrieval latency and storage while keeping relevance highly competitive.
More from Research
- Muennighoff presents infinite test-time scaling and prefix sliding work at MIT NLP Seminar — Muennighoff · 2026-10-11
- Symmetric cryptanalysis took 6,677 research-years, 6.4x all lattice work combined — matthew_d_green · 2026-10-11
- 10th grader used free Muse to produce 3 verified math preprints in 8 hours — EastConsequence3792 · 2026-10-11
- ML interatomic potentials reveal fast Li+ conduction mechanism in glassy antiperovskites — TimothyDuignan · 2026-10-11
- Bittensor subnet Metanova uses competition-driven AI drug discovery with robot labs — const_reborn · 2026-10-11
- Lean4 proofs are not a silver bullet: consistency gaps, soundness bugs, and flawed benchmarks — elie · 2026-10-11