Tencent releases universal multimodal embedding model, hits SOTA on MMEB
NielsRogge · x · 2026-08-25
Tencent has released a universal multimodal embedding model on Hugging Face that maps text, images, videos, and visual documents into a shared representation space, achieving SOTA on MMEB v2 and v3.
Per Niels Rogge, it also beats Meta's MoE vision model released just last week on COCO image-to-text retrieval, topping the Papers with Code leaderboard.
More from Models
- ChatGPT caught searching specific subreddits by name despite claims — gaganghotra_ · 2026-08-25
- LLM Memory Often Makes Things Worse — Maybe Forgetting Is the Optimal Process — sebpaquet · 2026-08-25
- Zhipu GLM 5.3 Flash interface potentially leaked online — LegacyRemaster · 2026-08-25
- Paper finds LLM skills vary by language; English reasoning recovers performance — LChoshen · 2026-08-25
- Dynamic quants of Qwen3.8 27B fix 'caveman thinking' issue — Glad_Claim_6287 · 2026-08-25
- 'What must be orchestrated today is model behavior tomorrow': voice pipelines get absorbed into models — morqon · 2026-08-25