WeMM-Embedding tops MMEB-v3: 9B scores 59.5, 2B beats every 7B/8B model

tomaarsen · x · 2026-09-02

On MMEB-v3 (190 tasks covering text, agent, and cross-modal retrieval), the WeMM-Embedding family scores:

The 2B variant already beats every 7B and 8B model listed.

The thread also shows a SentenceTransformer usage snippet with trustremotecode=True, encoding queries as text and documents as either text or {"video": ...} items.

Related event: Tencent Open-Sources WeMM-Embedding, Tops MMEB Benchmarks(6 posts)→

Original post →

More from Multimodal

Multimodal channel →