Understanding the modality gap in VLMs and contrastive learning mechanisms

gabriberton · x · 2026-08-19

Explains the modality gap in Vision-Language Models (VLMs). In a randomly initialized transformer, image embeddings cluster together. CLIP training attempts to pull positive pairs together and push negatives apart, but in practice, the pushing forces can be too strong, preventing the modality gap from closing.

Related event: Inside the Modality Gap in VLMs: Why It Persists and When It Helps(7 posts)→

Original post →

More from Multimodal

Multimodal channel →