Shared Geometry: TUM/MIT researchers align text and image spaces without paired data
kwangmoo_yi · x · 2026-10-10
Researchers from TU Munich, ETH Zurich and MIT (including Phillip Isola) present Shared Geometry as a Rosetta Stone, a cross-modal alignment method requiring no paired data: it aligns text and image embedding spaces through their internal geometric structure.
Key evaluation tools:
- FOSCTTM on 40,504 unseen COCO validation pairs (0 = true caption always nearest, 0.5 = random)
- CKA measuring structural similarity up to orthogonal transforms and scaling (1 = isomorphic)
- Gromov-Wasserstein matching using only within-space distances, no shared coordinates
- Orthogonal Procrustes for read-out
The headline result: the method never uses pairs during training, yet achieves cross-modal correspondence purely from each modality's internal geometry.
More from Multimodal
- Creator renders AI music video with local video model after 48 hours on two GPUs — tetsuoai · 2026-10-10
- Telling an LLM to "believe in yourself" helps it write 3D SDF models, but not enough — keenanisalive · 2026-10-10
- This 3D dragon is 27KB of LLM-generated GLSL, not a mesh or NeRF — keenanisalive · 2026-10-10
- OmniSeek turns Omni-LLMs into evidence-seeking agents, +15.5 points on VideoHolmes — mohitban47 · 2026-10-10
- Nikon Rescinds Microscopic Video Contest Win Over Generative AI Use — nordicinst · 2026-10-10
- AI fake videos have hit another level of realism, researcher warns — rohanpaul_ai · 2026-10-10