Ovis-Embedding Debuts: Unified Omni-Modal Embeddings Hit SOTA Across MMEB and More

Embedding Team · hf · 2026-09-23

The Embedding Team introduces Ovis-Embedding, an omni-modal embedding family that encodes text, image, video, and audio in a shared representation space instead of separate per-modality towers. Key advances: (1) native omni-modal initialization from a pretrained Qwen-omni backbone adapted via contrastive training with low-rank init; (2) data-centric training on a broad multimodal corpus with homogeneous-source sampling for task-consistent batches and informative in-batch negatives; (3) embedding-specific optimization using focal loss and similarity-based Embedding Distillation, plus low-rank feature decomposition for compact, flexible-dimensional embeddings at inference. It achieves state-of-the-art results on MMEB-v3, MMEB-v2, MVEB, MAEB, and RTEB, pointing toward universal any-to-any retrieval.

Related event: Ovis-Embedding Debuts: Unified Omni-Modal Embeddings Hit SOTA Across MMEB and More(2 posts)→

Original post →

More from Research

Research channel →