Linkup Open-Weights SparseUp: Strongest Sub-150M Sparse Retriever on BEIR
tomaarsen · x · 2026-09-17
Linkup released open-weight sparse embedding model SparseUp (Apache 2.0), praised by Sentence Transformers maintainer tomaarsen.
- Built on ModernBERT (149M), same backbone and fine-tuning data as LightOn's DenseOn/LateOn, completing the controlled three-architecture comparison
- Claims to be the strongest public vocabulary-based sparse encoder under 150M: 56.4 nDCG@10 on BEIR-13 (vs LateOn 57.9, DenseOn 56.9 under identical conditions)
- Three knobs vs vanilla SPLADE: logitshift=15 for sparse ReLU support, positiontopk=12 expansion budget per token, vocabfold collapsing case/space surface forms (vocab 50k → 34k)
- Contrastive fine-tuning on LightOn's mixture (7 hard negatives from 50 + in-batch), no cross-encoder distillation
Available on Hugging Face with inference-ready deployment.
Related event: Linkup Open-Sources SparseUp, Top Sparse Retriever Under 150M on BEIR-13(10 posts)→
More from Models
- OpenAI reportedly close to solving Hodge Conjecture, another $1M Millennium Prize problem — Distinct-Question-16 · 2026-09-18
- OpenAI reportedly close to solving Hodge Conjecture, another Millennium Prize problem — Distinct-Question-16 · 2026-09-18
- Sentdex Benchmarks LLMs on Halite: DSV4.1 Flash vs GLM 5.3 Flash in NVFP4 — Sentdex · 2026-09-18
- Six real-world uses for TypeSafe's Jev judgment model, 4x faster than Gemini in evals — HamelHusain · 2026-09-18
- Google DeepMind to host llama.cpp x Gemma side event at Nerdearla — mervenoyann · 2026-09-18
- Jina AI launches jina-ocr-v1, a 570M-active-param visual document parser — JinaAI_ · 2026-09-18