DINOv2 Meets Qwen3: Aligning Image and Text Embeddings With Zero Paired Data

phillip_isola · x · 2026-10-10

Phillip Isola's team (Dominik Schnaus et al., TU Munich/MIT/ETH) achieved a long-dreamt result: aligning image and text embedding spaces without a single image-caption pair. DINOv2 has never seen a caption and Qwen3 has never seen an image, yet aligning their shared geometry — even across different datasets — works. The method combines Gromov-Wasserstein matching and Orthogonal Procrustes, evaluated with FOSCTTM on 40,504 COCO pairs. A sign that representational geometry converges across modalities.

Related event: MIT Team Aligns Image and Text Embeddings Without Paired Data, Answering Platonic Representation Critiques(5 posts)→

Original post →

More from Research

Research channel →