DINOv2 and Qwen3 embeddings aligned without any image-text pairs

A new paper from TU Munich, MIT and ETH Zurich researchers shows that DINOv2, which never saw captions, and Qwen3, which never saw images, learn surprisingly similar world structures, and their embedding spaces can be aligned zero-shot without any image-text pairs.

2026-10-10 ~ 2026-10-10 · 4 related posts