Tencent Open-Sources WeMM-Embedding, a Top-Ranking Multimodal Model

abhishek__AI · x · 2026-08-28

Tencent's WeChat Vision Team has open-sourced WeMM-Embedding, a family of universal multimodal embedding models. Supporting unified representations for text, images, videos, and visual documents, it comes in 2B, 4B, and 9B sizes. The model achieves #1 on MMEB-v2 and v3 benchmarks, is designed for multimodal retrieval, and is 100% open source on Hugging Face.

Related event: Tencent Open-Sources WeMM-Embedding, Tops MMEB Leaderboard(3 posts)→

Original post →

More from Multimodal

Multimodal channel →