DINOv2 Meets Qwen3: Aligning Image and Text Embeddings With Zero Paired Data
phillip_isola · x · 2026-10-10
Phillip Isola's team (Dominik Schnaus et al., TU Munich/MIT/ETH) achieved a long-dreamt result: aligning image and text embedding spaces without a single image-caption pair. DINOv2 has never seen a caption and Qwen3 has never seen an image, yet aligning their shared geometry — even across different datasets — works. The method combines Gromov-Wasserstein matching and Orthogonal Procrustes, evaluated with FOSCTTM on 40,504 COCO pairs. A sign that representational geometry converges across modalities.
More from Research
- CS theorist Lance Fortnow on AI and math: don't panic, read a new proof — fortnow · 2026-10-10
- Skill Constellations: first dated copy network of 2.1M agent-skill adoptions on GitHub — UBC-O · 2026-10-10
- Physical attention bias cuts cable-simulation prediction error by 15%+ — Avihai Giuili · 2026-10-10
- GitSwarm paper: agents collaborate through a shared Git repo, solving all 30 IMOProofBench-Advanced problems — anirudhg9119 · 2026-10-10
- GitSwarm agents spontaneously adapt their workflow: building on coding tasks, verifying on proof tasks — anirudhg9119 · 2026-10-10
- GitSwarm's open-ended AutoResearch: agent swarms autonomously improve transformer architectures on GPUs — anirudhg9119 · 2026-10-10