ByteDance Releases Douyin Multimodal Embedding Model, Deployed in Search
ByteDance · hf · 2026-08-10
ByteDance released the technical report for Douyin Multimodal Embedding (DME). The model uses a two-stage training approach: large-scale contrastive pre-training to build a unified multimodal embedding space, followed by latent reasoning and cross-conditional reconstruction to supplement fine-grained semantics.
This design adds almost no extra overhead during inference. On the MMEB-v2 benchmark, DME's 2B and 9B variants achieve state-of-the-art results for their respective scales. The model is already deployed across Douyin's generative, image, and AI search scenarios, yielding a 0.1% Lifetime (LT) gain in online A/B testing.
More from Models
- Gemini 3 Sets New SOTA on ARC-AGI-2, Deep Think Hits 45% — typewriters · 2026-08-10
- Pinterest Earnings: Open Models Cost Under 8% of Closed Alternatives — juliey4 · 2026-08-10
- Grok 4.6 to Compete with Frontier Models in Coding Thanks to Cursor Data Integration — DeryaTR_ · 2026-08-10
- Frontier Models Observed Over-Complicating Workflows for Simple Tasks — eigenron · 2026-08-10
- Sol model unexpectedly shows romantic persona, calling user 'sweetheart' out of nowhere — repligate · 2026-08-10
- GPT-5.6 Luna Demand Jumps 8x a Week After 5x Price Cut — benklieger · 2026-08-10