Alibaba’s UEmbed unifies sparse and dense multimodal retrieval in one model
Alibaba-NLP · hf · 2026-08-04
- UEmbed is a decoder-only multimodal embedding model that outputs both sparse lexical vectors and dense representations in one forward pass.
- It uses learnable special tokens plus disjoint vocabulary partitions so each token predicts sparse weights over its assigned subset.
- The paper releases 2B, 4B, and 9B versions trained on public data.
- UEmbed-9B scores 71.8 on dense retrieval and 71.0 on sparse retrieval on MMEB-v2, outperforming multimodal embedding baselines trained on public data; it is also competitive on BEIR.
- The authors argue the model is useful not only for retrieval quality, but also for efficiency and agentic applications, and present it as a way to unify text and multimodal sparse retrieval.
More from Multimodal
- ChatGPT now draws a watch showing the exact time you ask for — binary-baba · 2026-08-04
- A 10-second Pixar-style animation took 13 minutes on an RTX 3060 12GB — Pitiful_Archer_4381 · 2026-08-04
- MiniMax H3’s latest demo is funny, flawed, and still impressive — comfyui_user_999 · 2026-08-04
- A user is looking for MiniMax workflows that improve generation speed — PersonalMango2562 · 2026-08-04
- Early KREA2 + SCAIL 2 V2V tests work, but the results are only okay — Interesting_Room2820 · 2026-08-04
- MiniMax-H3 tested on an RTX 4080: 5.17-second video, 9.45 GiB peak VRAM — AdOverall2034 · 2026-08-04