Why CLIP Training Fails to Close the Modality Gap in Multimodal Models
gabriberton · x · 2026-08-19
- The Phenomenon: In models like CLIP using separate transformers for image and text, embeddings from different modalities cluster in distinct regions, creating a "modality gap".
- The Intuition Trap: Despite contrastive learning pulling positive pairs together and pushing negatives apart, the expected closure of this gap does not occur in practice.
- Geometric Insight: In high-dimensional hyperspheres (1000 dims), the gap may exist in only a few dimensions. While cross-modal positive samples are close along most dimensions, the overall geometric separation persists.
Related event: Inside the Modality Gap in VLMs: Why It Persists and When It Helps(7 posts)→
More from Research
- Can public chat data predict real-world AI misalignments? — yoavartzi · 2026-08-20
- Research exposes LLM API vulnerability leaking hidden chain-of-thought — burkov · 2026-08-20
- The Human-or-Machine Issue: Turing-Inspired Reflections — ArtificialOther · 2026-08-20
- Meta Research Challenges Chinchilla Scaling Laws on Data-Compute Interactions — burkov · 2026-08-20
- 14,472 AI citations analyzed: business websites still win 60% of local search citations — gaganghotra_ · 2026-08-20
- LEGO-RL: harness-native reinforcement learning for coding agents — Lego-X · 2026-08-20