High-dimensional reality: Why cross-modal retrieval works despite the gap

gabriberton · x · 2026-08-19

Explains why cross-modal retrieval works despite the visible modality gap. The hypersphere is high-dimensional (usually 1000 dims), so the gap might exist only in one dimension. Along most dimensions, positive cross-modal samples are actually close, explaining the effectiveness of text-to-image retrieval.

Related event: Inside the Modality Gap in VLMs: Why It Persists and When It Helps(7 posts)→

Original post →

More from Multimodal

Multimodal channel →