Omni-Embed-Mini: A 0.9B Embedder Adds Five Modalities Without Touching Text Weights

_reachsumit · x · 2026-10-02

A new arXiv paper introduces Omni-Embed-Mini, a 0.9B-parameter omni-modal embedding model that maps text, speech, audio, images, video, and visually-rich documents into a single shared cosine space.

Original post →

More from Multimodal

Multimodal channel →