Ovis-Embedding Debuts: Unified Omni-Modal Embeddings Hit SOTA Across MMEB and More
Embedding Team · hf · 2026-09-23
The Embedding Team introduces Ovis-Embedding, an omni-modal embedding family that encodes text, image, video, and audio in a shared representation space instead of separate per-modality towers. Key advances: (1) native omni-modal initialization from a pretrained Qwen-omni backbone adapted via contrastive training with low-rank init; (2) data-centric training on a broad multimodal corpus with homogeneous-source sampling for task-consistent batches and informative in-batch negatives; (3) embedding-specific optimization using focal loss and similarity-based Embedding Distillation, plus low-rank feature decomposition for compact, flexible-dimensional embeddings at inference. It achieves state-of-the-art results on MMEB-v3, MMEB-v2, MVEB, MAEB, and RTEB, pointing toward universal any-to-any retrieval.
More from Research
- Google's Light Heads Cuts YouTube Recommender Experiment Cycles from Weeks to Days — _reachsumit · 2026-09-23
- Orthrus Serves Embedding and Generation in One GPU Batch, 4.52x RAG Throughput — _reachsumit · 2026-09-23
- Spotify: Behavioral Stats Boost LLM Reranking 13.3% but Teach It Shortcuts — _reachsumit · 2026-09-23
- CoVeR Cuts 62-68% of Agentic Retrieval Verifier Calls Without Losing Accuracy — _reachsumit · 2026-09-23
- Simons Institute–Jane Street Circles calls for small-group research proposals, deadline Oct. 15 — jasondeanlee · 2026-09-23
- jev-gc: reversible context garbage collection for long-running AI agents — Maleficent_College57 · 2026-09-23