FULL STORY
Tencent Open-Sources WeMM-Embedding, Tops MMEB Leaderboard
Tencent's WeChat vision team open-sourced the WeMM-Embedding family built on Qwen3.5, released a technical report the next day, with the 9B model topping the MMEB leaderboard.
2026-08-25 ~ 2026-08-28 · 2 episodes · 10 posts
Episode 1 · Tencent Open-Sources WeMM-Embedding Multimodal Models, 9B Tops MMEB v2/v3 (2026-08-25, 7 posts)
On August 25, Tencent released the WeMM-Embedding family on Hugging Face, followed by its technical report the next day. Built on Qwen3.5 and offered in 2B, 4B and 9B sizes, the models align text, images, video, visual documents and arbitrary interleaved multimodal inputs into a shared space, producing 4096-dim L2-normalized embeddings. The 9B variant tops MMEB v2 and v3, and the models are deployed at scale within WeChat.
Confirmed
- The family includes 2B, 4B and 9B variants based on Qwen3.5; the 2B version uses MRL (Multi-Resolution Representation Learning)
- Supports text, image, video, visual document and interleaved multimodal inputs with 4096-dim L2-normalized embeddings; audio is not yet supported
- The 9B model achieves SOTA on MMEB v2 and v3
- Per comparisons by @NielsRogge, the model leads on tasks such as COCO image-to-text retrieval
- Training proceeds in two stages: large-scale multimodal alignment followed by a curated-data stage
- Per @aigclink, the models were open-sourced by the WeChat vision team and already power search and recommendation in WeChat Channels, Official Accounts, Moments and e-commerce
- Per @reachsumit, even the 2B model surpasses 8B baselines
Why it matters
A unified multimodal embedding space is a key component for cross-modal retrieval and recommendation. WeMM-Embedding tops MMEB while remaining competitive at small scale, and its large-scale WeChat deployment provides real-world validation.
- Tencent releases WeMM-Embedding-2B multimodal embedding model — tencent · 2026-08-25
- Tencent releases WeMM-Embedding multimodal model family — jacek2023 · 2026-08-25
- Tencent releases universal multimodal embedding model, hits SOTA on MMEB — NielsRogge · 2026-08-25
- WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report — tencent · 2026-08-26
- Tencent releases WeMM-Embedding: 2B model beats 8B baselines — _reachsumit · 2026-08-26
- Tencent Open Sources WeMM-Embedding: 9B Model Ranks #1 on MMEB v2/v3 — aigclink · 2026-08-26
- Tencent open-sources WeMM multimodal embedding models, 9B ranks first on dual benchmarks — aigclink · 2026-08-26
Episode 2 · Tencent Open-Sources WeMM-Embedding, Tops MMEB Leaderboard (2026-08-27, 3 posts)
Tencent's WeChat vision team open-sourced the WeMM-Embedding family (2B/4B/9B), built on Qwen3.5, supporting unified text, image, video and visual document embeddings with MRL support. It topped the MMEB-v2 leaderboard and trended on Hugging Face.
- Tencent Open-Sources WeMM-Embedding-9B, a Multimodal Embedding Model Trending on HF — tencent · 2026-08-27
- Tencent Open-Sources WeMM-Embedding: A Top-Ranked Multimodal Model — abhishek__AI · 2026-08-28
- Tencent Open-Sources WeMM-Embedding, a Top-Ranking Multimodal Model — abhishek__AI · 2026-08-28