ByteDance Releases Douyin Multimodal Embedding Model, Deployed in Search

ByteDance · hf · 2026-08-10

ByteDance released the technical report for Douyin Multimodal Embedding (DME). The model uses a two-stage training approach: large-scale contrastive pre-training to build a unified multimodal embedding space, followed by latent reasoning and cross-conditional reconstruction to supplement fine-grained semantics.

This design adds almost no extra overhead during inference. On the MMEB-v2 benchmark, DME's 2B and 9B variants achieve state-of-the-art results for their respective scales. The model is already deployed across Douyin's generative, image, and AI search scenarios, yielding a 0.1% Lifetime (LT) gain in online A/B testing.

Original post →

More from Models

Models channel →