Adaptive Sparsity Optimization with Learnable Soft Top-K and Per-Term Thresholding for Efficient Retrieval
Wentai Xie, Parker Carlson, Shanxiu He, Tao Yang
cs.IR
2026-10-02
AdaSparse adds learnable soft top-K and per-term thresholds to Lion-SP. MRR@10 stays 0.408; query/doc length falls to 31%/26.6%; Seismic k=1000 latency drops from 9.0ms to 2.4ms.
Sparse retrieval learns a non-negative weight per vocabulary term. At query time it scans only the postings lists of non-zero terms, so a CPU can return candidates. SPLADE learns that expansion with BERT. Lion-SP swaps the encoder for LLaMA-3, reports leading relevance on MS MARCO and BEIR, and emits longer vectors.
LLaMA-3's vocabulary has 128,256 tokens, about 4.2 times BERT. One English word splits into case, number, and whitespace variants, and multilingual pretraining also lights up non-English tokens on English queries. On "how many years did william bradford serve as governor of plymouth colony?", SPLADE has length 20. On that same query Lion-SP-1B has length 272, with 9 English variants of "year" plus 2 Chinese tokens, and an average weight of 0.20 against 0.40 for SPLADE. Figure 1 shows MS MARCO non-zero scores piled up at low, similar values. Inverted-index latency and storage both grow with the non-zero count.
A fixed top-K keeps the same count on every vector and over-prunes complex queries. FLOPs regularization penalizes squared term weights inside a batch. Large noise shrinks, then the gradient fades on small weights and a long tail stays. Raising the coefficient costs relevance. Hybrid thresholding (HT) learns one global cutoff and can zero an entire vector out of domain. Mass-ratio pruning (MRP) keeps a head whose L1 mass hits a hand-set ratio alpha, so the surviving count is not learned. The introduction says official Lion-SP document vectors are 8.9 times one SPLADE variant. That variant's length is not given.
AdaSparse adds two regularizers to the Lion-SP loss. Inference does not cut at a hard top-K. Weights under a per-term threshold are zeroed, and the rest of the encoding matches Lion-SP.
STop is a soft cap that scales with the input, K = b times the original length. Defaults are b=10 for queries and b=8 for documents. Realized expansion can sit below b or a little above it. A sigmoid gate uses a pivot equal to the average of the K-th and (K+1)-th weights. Above the pivot the gate is near 0 and the penalty turns off. Below it the gate is near 1 and the squared weight is pulled down. Temperature gamma is 0.04.
Terms already present in the raw text are exempted with bias beta=2, above the non-zero weights typical in Figure 1, most of them under 1.4. In "quantum computing impact on finance", quantum scores 0.24. A global cutoff of 0.26 deletes it. STop keeps it because it occurs in the original text. A term sitting on the pivot with weight above 0.16 can get an upward gradient, so the ranking loss can lift a useful term back inside the soft cap. If the weight ranking is noisy, useful terms past the cap are penalized anyway.
PTT learns one threshold per token, with separate query and document sets. Training multiplies sub-threshold weights toward zero with a sigmoid of steepness pi=0.04, so gradients still flow. Indexing and query encoding hard-zero those weights. The threshold loss uses lambdaT=1. A term whose weights are flat across documents never gets a threshold high enough to delete it, and sparsity stalls.
FLOPs coefficients stay at Lion-SP's lambdaQ=0.05 and lambdaD=0.04, and STop uses the same pair. FLOPs shrinks large noisy activations. After a term falls under the pivot, STop keeps pushing, which is the region where the FLOPs gradient has already vanished. The ranking loss can push a useful term back up. The backbone is LLaMA-3.2 at 1B and 8B. Pretraining is masked next-token prediction on about 8.8 million MS MARCO passages, followed by one epoch of contrastive loss plus knowledge distillation. Loss comparisons start from the same pretrained weights. Raw queries average about 8 words and passages about 76.
In-domain numbers use 6,980 MS MARCO Dev queries and MRR@10, the reciprocal rank of the first relevant document in the top 10. TREC DL 2019 (43 judged queries), DL 2020 (53), and the 13 BEIR datasets use NDCG@10. The matched length comparison is a Lion-SP-1B run reproduced under the same training setup. The official checkpoint has query length 293, document length 1250, and BEIR 0.535. The reproduction has 208, 1052, and BEIR 0.545.
| Method | MRR@10 | DL19 | DL20 | BEIR | Query | Doc |
| Lion-SP-1B repro | 0.408 | 0.753 | 0.744 | 0.545 | 208 | 1052 |
| Heavier FLOPs | 0.400 | 0.745 | 0.743 | 0.530 | 130 | 315 |
| L1 + FLOPs | 0.404 | 0.754 | 0.718 | 0.531 | 98 | 297 |
| Fixed top-305 | 0.372 | 0.709 | 0.690 | 0.231 | 85 | 308 |
| VDR-256 | 0.398 | 0.741 | 0.717 | 0.493 | 103 | 272 |
| MRP alpha=0.75 | 0.406 | 0.748 | 0.742 | 0.529 | 78 | 268 |
| HT | 0.404 | 0.754 | 0.744 | 0.536 | 189 | 391 |
| AdaSparse-1B | 0.408 | 0.755 | 0.729 | 0.541 | 65 | 280 |
| Lion-SP-8B | 0.418 | 0.760 | 0.766 | 0.552 | 476 | 1457 |
| AdaSparse-8B | 0.417 | 0.763 | 0.752 | 0.543 | 85 | 297 |
AdaSparse-1B stays at Dev MRR@10 0.408. BEIR moves from 0.545 to 0.541, about 0.74% relative. Query length is 31% of the reproduction and document length is 26.6%. DL19 moves from 0.753 to 0.755. DL20 falls from 0.744 to 0.729. Fixed top-305 drops BEIR to 0.231. MRP lands at BEIR 0.529, about 2.3% lower, at a similar length. HT still has query length 189 and document length 391, about 40% longer documents than AdaSparse's 280, with Dev MRR 0.404.
STop alone, on top of FLOPs, cuts queries from 208 to 83 and documents from 1052 to 364, with MRR 0.409 and BEIR 0.540. Adding PTT reaches 65 and 280, MRR 0.408, BEIR 0.541. Swapping PTT for one global threshold gets queries to 46 and BEIR down to 0.530.
Seismic is measured on an AMD EPYC 9R45 with 256GB of RAM, three-run average, approximate search only. At depth 1000, AdaSparse-1B takes 2.390 ms against 9.048 ms for official Lion-SP-1B, about 3.8 times. At depth 10 the times are 0.292 ms and 1.179 ms, about 4 times. The index is 33GB against 138GB, about 4.2 times. HT takes 3.223 ms and 58GB. AdaSparse-8B uses 34GB and 2.863 ms at depth 1000.
On the 13 BEIR sets, official Lion-SP-1B averages query length 649, document length 1811, and NDCG 0.535. AdaSparse averages 129, 456, and 0.541. On MS MARCO the realized expansion is 8.1 times for queries and 3.7 times for documents, both under the training value of b. Query expansion stays within 10 times on 11 of 14 sets. NFCorpus, SciFact, and FEVER reach 11.7 to 13.5 times. With query b=6 and document b=5, lengths are 50 and 206, Dev MRR is 0.409, and BEIR is 0.537. At a similar document length, heavier FLOPs scores 0.382 / 0.508 and MRP scores 0.394 / 0.526.
SPLADEv3, BERT at 110M, has query length 24, document length 170, Dev MRR 0.402, and BEIR 0.517. AdaSparse-1B queries are still about 2.7 times longer and documents about 1.6 times longer. Dense RepLLaMA (LLaMA-2 7B) uses an index of about 145GB. AdaSparse-8B uses 34GB, about 4.3 times smaller, with Dev 0.417 against 0.412, DL19/DL20 at 0.763/0.752 against 0.743/0.725, and BEIR 0.543 against 0.551. Scaling 1B to 8B lifts Dev from 0.408 to 0.417 and BEIR from 0.541 to 0.543, while query length grows from 65 to 85 and document length from 280 to 297.
Much of what the 128K vocabulary adds is inflection variants and cross-lingual neighbors. A system that already runs Lion-SP can retrain this loss and re-encode documents. Dot-product scoring and the inverted index stay as they are. Seismic latency follows document length. Growing queries from length 9 to 69 barely moves it, which is why document expansion sits at 3.7 times and query expansion at 8.1 times on MS MARCO. That split is tied to Seismic. Other inverted-index engines are not measured.
Against SPLADEv3 the vectors are still longer and in-domain relevance is higher. Against dense RepLLaMA the index is about 4.3 times smaller and average BEIR is about 1.5% lower. When the budget is CPU inverted-index search, that trade holds.
STop trusts the model's ranking of term weights. The paper calls the failure a brittle cliff, mitigates it with a larger b, and never measures how much useful expansion the noise removes. PTT stalls on flat per-term distributions: raising lambdaT from 1 to 6 drops BEIR from 0.541 to 0.534 while length only moves from 65/280 to 54/245. beta=2 and gamma=0.04 are taken from Lion-SP's weight histogram on MS MARCO, and the method is trained and tested only on that backbone. Empty vectors are named as a risk of a global threshold. Whether AdaSparse itself emits empty queries or documents on BEIR is not reported.
Averages hide set-level moves. Against the reproduction, DL20 falls from 0.744 to 0.729, and Lion-SP-8B falls from 0.766 to 0.752. Against official Lion-SP-1B, ArguAna falls from 0.488 to 0.470 and Climate-FEVER from 0.304 to 0.281, while Quora rises from 0.791 to 0.858 and TREC-COVID from 0.785 to 0.817.
The speed comparison uses official Lion-SP, exact MRR 0.410. The shorter reproduction has no Seismic timing, so the 4.2 times storage gap mixes in a checkpoint that was already denser. Against the reproduction, length only falls to 31% and 26.6%. Query cutoffs differ too: 8 for HT and 6 for AdaSparse. At depth 10, approximate MRR for AdaSparse-1B is 0.401, 0.007 under the exact 0.408, outside the 0.002 band stated in the paper. At depth 1000 the gap is 0.002. Training also skips MS MARCO title markers. The Lion row in Table 13 is the 8B model, query 85 and document 297. The 2.7 times and 1.6 times figures are the 1B model against SPLADEv3.