Aligning DINOv2 and Qwen3 embedding spaces with zero image-caption pairs

AccBalanced · x · 2026-10-11

dominikschnaus presents research aligning the embedding spaces of DINOv2 (which has never seen a caption) and Qwen3 (which has never seen an image) without a single image-caption pair — it even works when images and captions come from different datasets. Project page and paper are available.

Related event: Shared Geometry as a Rosetta Stone: DINOv2 and Qwen3 Embeddings Aligned Without Any Image-Text Pairs(7 posts)→

Original post →

More from Multimodal

Multimodal channel →