Image-only DINOv2 and text-only Qwen3 converge on similar world geometry, new paper finds
lmoroney · x · 2026-10-10
A new paper, "Shared Geometry as a Rosetta Stone" by Dominik Schnaus et al. (TU Munich, MIT, ETH Zurich), shows that DINOv2 (trained only on images) and Qwen3 (trained only on text) have surprisingly similar internal representational geometry.
- Method: aligning the two embedding spaces using unpaired COCO data — images from one half, captions from the other — via Gromov-Wasserstein matching plus Orthogonal Procrustes, achieving cross-modal alignment without paired data.
- Evaluation: FOSCTTM retrieval on 40,504 COCO validation pairs (random guessing = 0.5), and CKA to compare sample arrangement across spaces (invariant to orthogonal transforms and scaling).
- Takeaway: vision and language models converge on shared structure despite disjoint training modalities, supporting "shared geometry" as a cross-modal Rosetta Stone. Includes an interactive visualization.
Related event: DINOv2 and Qwen3 embeddings aligned without any image-text pairs(4 posts)→
More from Models
- An AI answered with just a thumbs-up instead of a verbose reply — users love it — enggirlfriend · 2026-10-10
- Blind test: Astra beats Opus 5.5 80% of the time on hardcore coding — bindureddy · 2026-10-10
- Liquid AI's decision model d1 lands on Vercel AI Gateway with vision support — maximelabonne · 2026-10-10
- Open TTS Leaderboard adds Paradee-8M, a Kokoro-82M distill matching WER at 1/10 params — realmrfakename · 2026-10-10
- Kimi gateway latency test ranks GitHub first, Neon second, ngrok third — mariorod1 · 2026-10-10
- Grok in group chats is 'really nice', but users note messages are no longer private — Angaisb_ · 2026-10-10