Alibaba open-sources 3 Ovis multimodal embedding models under Apache 2.0
AdinaYakup · x · 2026-09-23
Alibaba's ATH MaaS released three Ovis universal multimodal embedding models, all Apache 2.0 licensed and live on Hugging Face:
- Omni-3B: built on Qwen 2.5 Omni, supports audio input, self-reported MMEB-v3 58.46 (+5.19 over best baseline)
- VL-2B: based on Qwen 3.5 2B, best value, MMEB-v2 77.46
- VL-9B: based on Qwen 3.5 9B, highest accuracy at MMEB-v2 81.13
An accompanying paper (arXiv:2609.25165) is available; all scores are self-reported.
Related event: Alibaba Open-Sources Ovis-Embedding, Setting New Multimodal SOTAs(4 posts)→
More from Models
- GPT-6 Astra beats Claude Opus 5.5 31s in LLM-driven robot sumo sim — DJiafei · 2026-09-23
- JevBench v1.4 follow-up link: details of the anti-benchmaxxing methodology — airesearch12 · 2026-09-23
- Casually mentioning a cat makes Opus 5.5 grill the user for cat facts — voooooogel · 2026-09-23
- JevBench v1.4 adds 308 evolving sealed tasks to stop benchmaxxing, now covers 70+ models — airesearch12 · 2026-09-23
- Stop benchmarking LLMs with 3D games, says Abacus.AI CEO — labs fine-tune for it — bindureddy · 2026-09-23
- Opus 5.5 impresses with a hand-crafted Mona Lisa SVG — djcows · 2026-09-23