Tencent releases universal multimodal embedding model, hits SOTA on MMEB

NielsRogge · x · 2026-08-25

Tencent has released a universal multimodal embedding model on Hugging Face that maps text, images, videos, and visual documents into a shared representation space, achieving SOTA on MMEB v2 and v3.

Per Niels Rogge, it also beats Meta's MoE vision model released just last week on COCO image-to-text retrieval, topping the Papers with Code leaderboard.

Original post →

More from Models

Models channel →