ConCor-1 Reshapes Vision-Language Grounding with 29% Zero-Shot F1 Gain
raivn · hf · 2026-08-11
Existing vision-language grounding is typically reduced to a unidirectional localization problem (finding image regions for predefined text), overlooking the basic challenge of determining which parts of the text are visually referential.
The paper reformulates grounding as bidirectional concept correspondence over an image-text pair: recovering all correspondences between visually referential text spans and instance-level image segments without assuming relevant text spans are provided. This unifies common grounding tasks like phrase grounding, referring expression grounding, and open-vocabulary detection.
To address this, researchers introduced ConCor-1, built on top of a pretrained vision-language model. It uses learnable bridge tokens to represent candidate image-text correspondences, predicting a text mask, an image mask, and a correspondence presence score. Experiments show that ConCor-1 consistently outperforms baselines, improving correspondence F1 by 48% on a long-caption dataset and by 29% on zero-shot LVIS.
More from Research
- Master LLMs from Scratch in 60 Days: 8 Essential Papers — thisguyknowsai · 2026-08-11
- Researchers Expose API Flaw: Encrypted Chain-of-Thought in Major LLMs Can Be Stolen — dpaleka · 2026-08-11
- DeepSeek Vision Paper: Integrating Spatial Markers into Reasoning Trajectories — teortaxesTex · 2026-08-11
- DeepSeek's Recent Two Major Releases Leave Next Paper Direction a Mystery — teortaxesTex · 2026-08-11
- Luth-2 Released: Sets New SOTA for French Small Language Models — Unusual_Shoe2671 · 2026-08-11
- Open-Source SynthID-Text-Detector: Reference Implementation for AI Text Watermarking — jedisct1 · 2026-08-11