DeBias-CLIP tackles CLIP’s long-caption bias and hits state-of-the-art retrieval

Mila_Quebec · x · 2026-07-21

Vision-language models such as CLIP are biased toward the first words in long captions. The post says the team trained DeBias-CLIP to remove that shortcut and reports state-of-the-art results on long-text retrieval.

Original post →

More from Research

Research channel →