DeBias-CLIP tackles CLIP’s long-caption bias and hits state-of-the-art retrieval
Mila_Quebec · x · 2026-07-21
Vision-language models such as CLIP are biased toward the first words in long captions. The post says the team trained DeBias-CLIP to remove that shortcut and reports state-of-the-art results on long-text retrieval.
- The core issue is a positional bias in long-caption training.
- DeBias-CLIP is presented as the fix.
- The attached comparison image shows CLIP and Long-CLIP failing, while DeBias-CLIP succeeds.
More from Research
- NVIDIA pitches World Foundation Models as a way to scale physical AI data generation — MonaJalal_ · 2026-07-22
- Sol-Engine Boosts Video Generation Speed by up to 5x with Training-Free Sparse Attention — songhan_mit · 2026-07-22
- ICML 2026 paper CPMöbius lets models generate their own reasoning tasks — Dr_Singularity · 2026-07-22
- MIT's New Algorithm Speeds Up AI Diffusion Sampling, Wins ICML Outstanding Paper — MIT_CSAIL · 2026-07-22
- NeurIPS 2026 Workshop on Geometric Deep Learning and Optimal Transport Opens Call for Papers — ninamiolane · 2026-07-22
- RoboMME Podcast Preview: Benchmarking Memory for Robotic Policies — chris_j_paxton · 2026-07-21