Tencent releases WeMM-Embedding: 2B model beats 8B baselines
_reachsumit · x · 2026-08-26
Tencent released the WeMM-Embedding technical report, a family of universal multimodal embedding models supporting text, images, videos, visual documents, and arbitrarily interleaved inputs. The family includes 2B, 4B, and 9B variants. Training involves two stages: large-scale multimodal alignment followed by refinement using curated data, fine-grained relevance supervision, and cross-scale knowledge transfer. Evaluations show leading performance on public benchmarks: the 2B variant surpasses the previous 8B open-source baseline on MMEB-v2, while the 9B variant achieves a new state-of-the-art overall score of 80.6. The model has been deployed at scale in WeChat applications, including Channels, Official Accounts, Moments, and e-commerce services, yielding substantial gains.
Related event: Tencent Releases WeMM-Embedding Family of Multimodal Embedding Models(5 posts)→
More from Multimodal
- Emotional animation of girl touching sky whale generated by Google Gemini — michaelrabone · 2026-08-26
- AI generated dance video shows cool moves — No-Bookkeeper-char · 2026-08-26
- Face Anything: 4D Face Reconstruction from Any Image Sequence (ECCV 2026) — rsasaki0109 · 2026-08-26
- Ref2V demo: H3 model handles complex prompts in just 15 seconds — Jeffu · 2026-08-26
- GPT Image 2 technique transforms everyday photos into shape translation posters — RealJamesOfficial · 2026-08-26
- Pavo launches AgnesVideo 2.5 with free tier to cut AI short drama costs — 新智元 · 2026-08-26