DINOv2 and Qwen3 align with a single rotation matrix — no paired data needed
lmoroney · x · 2026-10-10
A new paper by Dominik Schnaus et al. (TU Munich, MIT and partners) shows image-only and text-only models converge on surprisingly similar world representations.
- Using DINOv2 (images only) and Qwen3 (text only), trained on images from one half of MS COCO and captions from the other — zero paired examples;
- The method learns a single orthogonal map (rotation/reflection only) that moves image embeddings into caption space, landing each picture near captions about the same content;
- It still works across different datasets; alignment is coarse and model-pair dependent, and CKA similarity on a small paired sample predicts how well it will go.
The interactive project page doubles as a lovely classroom demo for teaching embeddings.
Related event: DINOv2 and Qwen3 embeddings aligned without any image-text pairs(4 posts)→
More from Research
- Closed-loop experiments across 25 neural sites reveal hidden differences in model-brain alignment — ShahabBakht · 2026-10-10
- acceptodds builds a prediction market where researchers bet on ICLR 2027 paper acceptances — prof_kamilov · 2026-10-10
- Surge AI launches sudo L7: a benchmark testing whether coding agents can act like staff engineers — rmcwhorter99 · 2026-10-10
- Live feed shows what images AI agents use while hunting for new planets — BLUECOW009 · 2026-10-10
- Jeremy Avigad's slides on the future of mathematics in the age of AI — ChengleiSi · 2026-10-10
- Tetris RL experiment: pretraining caps what RL can reach — PPO can provably converge to a bad policy — shizhediao · 2026-10-10