DINOv2 and Qwen3 embeddings aligned without any image-text pairs
A new paper from TU Munich, MIT and ETH Zurich researchers shows that DINOv2, which never saw captions, and Qwen3, which never saw images, learn surprisingly similar world structures, and their embedding spaces can be aligned zero-shot without any image-text pairs.
2026-10-10 ~ 2026-10-10 · 4 related posts
- DINOv2 and Qwen3 embedding spaces aligned with zero image-caption pairs — TimDarcet · 2026-10-10
- DINOv2 and Qwen3 Embedding Spaces Aligned Without a Single Image-Caption Pair — NandoDF · 2026-10-10
- DINOv2 and Qwen3 align with a single rotation matrix — no paired data needed — lmoroney · 2026-10-10
- Image-only DINOv2 and text-only Qwen3 converge on similar world geometry, new paper finds — lmoroney · 2026-10-10