Tencent open-sources WeMM-Embedding: unified text/image/video embeddings, Apache 2.0

tomaarsen · x · 2026-09-02

Tencent released WeMM-Embedding, a family of universal multimodal embedding models that embed text, images, videos, and visual documents into a single vector space for cross-modal retrieval.

All three sizes (2B, 4B, 9B) are built on Qwen3.5 and released under Apache 2.0, now available on Hugging Face.

Related event: Tencent Open-Sources WeMM-Embedding, Tops MMEB Benchmarks(6 posts)→

Original post →

More from Models

Models channel →