DINOv2 and Qwen3 Embedding Spaces Aligned Without a Single Image-Caption Pair
NandoDF · x · 2026-10-10
Researchers aligned the embedding spaces of DINOv2 (never trained on captions) and Qwen3 (never trained on images) without using any image-caption pairs, bridging the semantic gap between language and perceptual representations. The alignment even holds when images and captions come from different datasets. Retweeted by Sander Dieleman, who calls it an exciting result he may expand on in a blog post.
Related event: DINOv2 and Qwen3 embeddings aligned without any image-text pairs(4 posts)→
More from Research
- Gym-Anything wins best paper at COLM 2026 Lifelong Agents Workshop — dan_fried · 2026-10-10
- Ai2's Olmo Hybrid hits Olmo 3 7B MMLU accuracy with 49% fewer training tokens — allen_ai · 2026-10-10
- SF event Oct 22 asks what it takes for AI to run scientific research end to end — nlarusstone · 2026-10-10
- Paper argues nested von Neumann architecture can make a million processors act as one computer — bronzeagepapi · 2026-10-10
- UPenn paper unifies diffusion and autoregression on one corruption lattice to predict decoding costs — upenn · 2026-10-10
- CMU's Keenan Crane wraps AI-assisted 3D modeling thread, releases all model files under CC0 — keenanisalive · 2026-10-10