Tencent releases WeMM-Embedding multimodal models for mixed text, image, and video retrieval

tomaarsen · x · 2026-09-02

Tencent's team released the WeMM-Embedding family (2B/4B/9B) on Hugging Face, aimed at retrieval over mixed media, especially with video in the corpus:

A technical report (arXiv 2608.24053) is out; the 2B variant has passed 7k downloads.

Related event: Tencent Open-Sources WeMM-Embedding, Tops MMEB Benchmarks(6 posts)→

Original post →

More from Models

Models channel →