ConCor-1: Reimagining vision-language grounding as bidirectional concept correspondence

RanjayKrishna · x · 2026-08-14

Researchers introduce ConCor-1, revisiting vision-language grounding. Traditional methods are unidirectional: language guides the model to locate objects in images. ConCor-1 formulates grounding as bidirectional concept correspondence, where the model jointly discovers referential concepts in language and vision and explicitly recovers their correspondence without pre-specified text spans. The authors believe this moves grounding from answering 'where is this?' to understanding the full correspondence structure between language and vision. ConCor-1 is the first step; the team is already extending this direction.

Related event: ConCor-1 Reframes Vision-Language Grounding as Bidirectional Concept Correspondence(2 posts)→

Original post →

More from Multimodal

Multimodal channel →