DINOv2 and Qwen3 embedding spaces aligned with zero image-caption pairs

DocXavi · x · 2026-10-11

New work extends ASIF (arXiv:2210.01738): DINOv2 and Qwen3 embeddings are aligned without a single image-caption pair, even across different datasets — questioning the need for massive paired multimodal training.

Related event: Shared Geometry as a Rosetta Stone: DINOv2 and Qwen3 Embeddings Aligned Without Any Image-Text Pairs(7 posts)→

Original post →

More from Research

Research channel →