Vision-Language Grounding Reformulated as Bidirectional Concept Correspondence

RanjayKrishna · x · 2026-08-14

Traditional vision-language grounding is often reduced to a unidirectional localization problem (finding image regions for given text), overlooking the challenge of determining which parts of the text are visually referential.

The paper Vision-Language Grounding as Bidirectional Concept Correspondence formulates grounding as recovering all correspondences between visually referential text spans and instance-level image segments without pre-provided text. This unifies tasks like phrase grounding, referring expression grounding, and open-vocabulary detection.

The authors introduce ConCor-1, a model built on a pretrained VLM. It uses learnable bridge tokens to represent candidate image-text correspondences, predicting a text mask, an image mask, and a correspondence presence score for each token.

Related event: ConCor-1 Reframes Vision-Language Grounding as Bidirectional Concept Correspondence(2 posts)→

Original post →

More from Research

Research channel →