Alibaba's Ovis-Embedding tops five benchmarks with unified omni-modal embeddings
_reachsumit · x · 2026-09-23
Alibaba introduced Ovis-Embedding, an omni-modal embedding family encoding text, image, video and audio in one shared space via a pretrained Qwen-omni backbone adapted with contrastive training. Key techniques include homogeneous-source sampling for task-consistent batches, focal loss, similarity-based embedding distillation, and low-rank decomposition for flexible embedding dimensions. It achieves SOTA on MMEB-v3, MMEB-v2, MVEB, MAEB and RTEB; paper and code are public.
More from Models
- Grok 4.7 flops in 100 multi-agent coding evals despite insightful solutions — teortaxesTex · 2026-09-23
- Will rumored GPT-6 'Sol' actually ship inside ChatGPT? — flowersslop · 2026-09-23
- Claude Opus 5.5 Said to Fall Back on Frontier LLM Dev Tasks, Drawing Fire — basedjensen · 2026-09-23
- Higgsfield demo: Claude Opus 5.5 crushes GPT-6 Astra at 3D game generation — VraserX · 2026-09-23
- User Gives Opus 5.5 Creative Tools and Asks What It Dreams About — angrypenguinPNG · 2026-09-23
- Yuchen Jin: Opus 5.5 underwhelms, frontier LLM coding has plateaued — Yuchenj_UW · 2026-09-23