DINOv2 and Qwen3 Embedding Spaces Aligned Without a Single Image-Caption Pair

NandoDF · x · 2026-10-10

Researchers aligned the embedding spaces of DINOv2 (never trained on captions) and Qwen3 (never trained on images) without using any image-caption pairs, bridging the semantic gap between language and perceptual representations. The alignment even holds when images and captions come from different datasets. Retweeted by Sander Dieleman, who calls it an exciting result he may expand on in a blog post.

Related event: DINOv2 and Qwen3 embeddings aligned without any image-text pairs(4 posts)→

Original post →

More from Research

Research channel →