Why CLIP training fails to close the modality gap in multimodal models

gabriberton · x · 2026-08-19

Explores why the modality gap persists. In CLIP's dual-transformer setup, the pushing forces on negative pairs are often too strong (each sample pushes away all others except its pair), preventing the gap from closing. This phenomenon, studied in the 'Mind the Gap' paper, occurs across many multimodal models.

Related event: Inside the Modality Gap in VLMs: Why It Persists and When It Helps(7 posts)→

Original post →

More from Multimodal

Multimodal channel →