Tencent releases WeMM-Embedding: 2B model beats 8B baselines

_reachsumit · x · 2026-08-26

Tencent released the WeMM-Embedding technical report, a family of universal multimodal embedding models supporting text, images, videos, visual documents, and arbitrarily interleaved inputs. The family includes 2B, 4B, and 9B variants. Training involves two stages: large-scale multimodal alignment followed by refinement using curated data, fine-grained relevance supervision, and cross-scale knowledge transfer. Evaluations show leading performance on public benchmarks: the 2B variant surpasses the previous 8B open-source baseline on MMEB-v2, while the 9B variant achieves a new state-of-the-art overall score of 80.6. The model has been deployed at scale in WeChat applications, including Channels, Official Accounts, Moments, and e-commerce services, yielding substantial gains.

Related event: Tencent Releases WeMM-Embedding Family of Multimodal Embedding Models(5 posts)→

Original post →

More from Multimodal

Multimodal channel →