DINOv2 and Qwen3 align with a single rotation matrix — no paired data needed

lmoroney · x · 2026-10-10

A new paper by Dominik Schnaus et al. (TU Munich, MIT and partners) shows image-only and text-only models converge on surprisingly similar world representations.

The interactive project page doubles as a lovely classroom demo for teaching embeddings.

Related event: DINOv2 and Qwen3 embeddings aligned without any image-text pairs(4 posts)→

Original post →

More from Research

Research channel →