Vision-Language Grounding Reformulated as Bidirectional Concept Correspondence
RanjayKrishna · x · 2026-08-14
Traditional vision-language grounding is often reduced to a unidirectional localization problem (finding image regions for given text), overlooking the challenge of determining which parts of the text are visually referential.
The paper Vision-Language Grounding as Bidirectional Concept Correspondence formulates grounding as recovering all correspondences between visually referential text spans and instance-level image segments without pre-provided text. This unifies tasks like phrase grounding, referring expression grounding, and open-vocabulary detection.
The authors introduce ConCor-1, a model built on a pretrained VLM. It uses learnable bridge tokens to represent candidate image-text correspondences, predicting a text mask, an image mask, and a correspondence presence score for each token.
More from Research
- MoME: Context-Aware Mixture-of-Memory Embeddings Outperform Value Embedding Baselines at Iso-FLOPs — vector-institute · 2026-09-21
- FRAUDSkill Boosts Audio Anti-Fraud Macro-F1 to 73.5% With Frozen Weights — PPSUCTeleantifraudCommunity · 2026-09-21
- TeleAntiFraud 2.0: Chinese Audio Fraud Benchmark Shows F1 Drops to 0.65 on Near-Domain Negatives — PPSUCTeleantifraudCommunity · 2026-09-21
- Fitting a performance cone with bootstrapped models to test if AI leaderboard rankings actually hold — PTenigma · 2026-09-21
- Researchers pitch peer review fixes: dedicated screeners, sanctions for bad submissions, review ratios — m2saxon · 2026-09-21
- AI writing detectors flagged her lab vision post; she asks what we should actually measure — furongh · 2026-09-21