Alibaba's Ovis-Embedding tops five benchmarks with unified omni-modal embeddings

_reachsumit · x · 2026-09-23

Alibaba introduced Ovis-Embedding, an omni-modal embedding family encoding text, image, video and audio in one shared space via a pretrained Qwen-omni backbone adapted with contrastive training. Key techniques include homogeneous-source sampling for task-consistent batches, focal loss, similarity-based embedding distillation, and low-rank decomposition for flexible embedding dimensions. It achieves SOTA on MMEB-v3, MMEB-v2, MVEB, MAEB and RTEB; paper and code are public.

Related event: Alibaba Releases Ovis-Embedding, Unifying All Modalities in One Vector Space(2 posts)→

Original post →

More from Models

Models channel →