Omni-Embed-Mini: A 0.9B Embedder Adds Five Modalities Without Touching Text Weights
_reachsumit · x · 2026-10-02
A new arXiv paper introduces Omni-Embed-Mini, a 0.9B-parameter omni-modal embedding model that maps text, speech, audio, images, video, and visually-rich documents into a single shared cosine space.
- Key insight: the text encoder stays completely frozen (bit-identical weights), so text retrieval cannot regress. The teacher signal is simply the frozen backbone's embedding of a dense cascaded caption paired with each media sample — since teacher and student share the same backbone geometry, lightweight projectors plus phased LoRA adapters on the modality encoders suffice for alignment.
- Training recipe: a Matryoshka SigLIP contrastive loss combined with an online hybrid hard-negative miner whose negatives sharpen as the encoder improves.
- Results: 49.57 nDCG@10 on MTEB-v2 BEIR-8 for text retrieval, while being 2.7x–9.5x smaller than compared open omni embedders; the recipe scales to a 2.3B variant by swapping in a native vision-language backbone.
More from Multimodal
- PEARL Debuts First User-History Personalized Image Generation Benchmark, +15% Over Baselines — Bo Ni · 2026-10-02
- First systematic survey of joint video-audio generation and editing: 9 categories, 28 edit types — Abhinav Sharma · 2026-10-02
- Full AI music video for NCT 127 built with Codex, Higgsfield MCP and ComfyUI — leesysysysy · 2026-10-02
- Image editing demo with Nano Banana — tkasasagi · 2026-10-02
- Creator: Opus 5.5 edits videos well, but forcing AI to clip without real need yields garbage — AlchainHust · 2026-10-02
- Reddit user's VEC concept car AI video shows startlingly realistic motion — Vashukanni · 2026-10-02