Sparse Embeddings Explained: 99% Zeros, Cheap to Store, Human-Readable
tomaarsen · x · 2026-09-17
tomaarsen, maintainer of Sentence Transformers, breaks down sparse embedding models (SPLADE-style):
- They embed text into vectors with vocabulary-size dimensionality, 99% zeros, only a few active dims
- Sparsity makes them cheap to store (save only active dims) and interpretable: an active dim 7429 corresponds to vocab token 7429, so embeddings are human-understandable
- Typically trained for retrieval, and with strong indexes they're extremely fast
Part of a thread discussing Linkup's newly open-weighted SparseUp sparse embedding model.
Related event: Linkup Open-Sources SparseUp, Top Sparse Retriever Under 150M on BEIR-13(10 posts)→
More from Research
- OpenAI reportedly close to solving Hodge Conjecture, another $1M Millennium Prize problem — Distinct-Question-16 · 2026-09-18
- OpenAI reportedly close to solving Hodge Conjecture, another Millennium Prize problem — Distinct-Question-16 · 2026-09-18
- Diffusion-augmented LLM Uno delivers lossless 2.2x speedup over autoregressive generation — HongyiWang10 · 2026-09-18
- Caltech trio on how AI is transforming scientific discovery, from quantum to ecology — AnimaAnandkumar · 2026-09-18
- Porting Agents' Last Exam's Linux CLI subset surfaced benchmark defects, yielding ALE-Gold — dejavucoder · 2026-09-18
- AISTATS 2027 opens submissions with AI review as a first-time feature — qberthet · 2026-09-18