Tencent releases WeMM-Embedding multimodal models for mixed text, image, and video retrieval
tomaarsen · x · 2026-09-02
Tencent's team released the WeMM-Embedding family (2B/4B/9B) on Hugging Face, aimed at retrieval over mixed media, especially with video in the corpus:
- Interleaved inputs: multiple images or videos in a single input; documents can mix {"video": ..., "text": ...}.
- Matryoshka truncation for smaller/faster embeddings.
- Usage: SentenceTransformer("tencent/WeMM-Embedding-9B", trustremotecode=True) with encodequery / encodedocument.
A technical report (arXiv 2608.24053) is out; the 2B variant has passed 7k downloads.
Related event: Tencent Open-Sources WeMM-Embedding, Tops MMEB Benchmarks(6 posts)→
More from Models
- Rumor: two major open-source model releases expected in September — lqiao · 2026-09-02
- Qwen3.8-Max-0902 debuts at #1 on Code Arena WebDev with 1691 pts, beating Claude Opus 5 — theimposingshadow · 2026-09-02
- Bug Hunt Bench: Fable 5.1 Low Beats Opus 5 Max at Lower Cost — PawelHuryn · 2026-09-02
- Gemini 3.8 Flash spotted in GCP Agent Studio, aimed at multimodal and coding tasks — testingcatalog · 2026-09-02
- Reported 90% on ARC-AGI-2 at $3.12/task with 32% cost reduction — eyishazyer · 2026-09-02
- Users notice GPT now starts ~80% of answers with "Yes" — Standard-Metal-3836 · 2026-09-02