Shared Geometry: TUM/MIT researchers align text and image spaces without paired data

kwangmoo_yi · x · 2026-10-10

Researchers from TU Munich, ETH Zurich and MIT (including Phillip Isola) present Shared Geometry as a Rosetta Stone, a cross-modal alignment method requiring no paired data: it aligns text and image embedding spaces through their internal geometric structure.

Key evaluation tools:

The headline result: the method never uses pairs during training, yet achieves cross-modal correspondence purely from each modality's internal geometry.

Related event: MIT Team Aligns Image and Text Embeddings Without Paired Data, Answering PRH Critics(6 posts)→

Original post →

More from Multimodal

Multimodal channel →