Vision-OPD uses 6.2K synthetic samples to boost fine-grained vision understanding
小红书技术REDtech · wechat · 2026-07-28
Xiaohongshu’s dots team introduces Vision-OPD, a vision-language post-training method designed to close the “region-to-global gap” in fine-grained visual understanding.
The paper argues that multimodal models often answer correctly when shown a cropped, zoomed-in region but fail on the full image. Vision-OPD turns that advantage into training signal by using an online self-distillation setup: a cropped-image teacher and a full-image student share the same model, while token-level divergence is minimized on the student’s own sampled trajectories.
Key results:
- Only about 6.2K fully synthetic samples were used.
- A 9B model reached 79.7 average accuracy across fine-grained benchmarks and ranked first overall.
- A 4B version reached 77.1 and also beat larger closed and open models.
- Performance on broader visual tasks stayed intact, suggesting the method improves fine detail perception without forgetting general vision skills.
More from Multimodal
- AI Rebuilds Homer's Odyssey into a 135-Minute Feature Film — heypearlai · 2026-07-28
- Row-Bot demo turns rough launch notes into decks, posts, and design assets — Acceptable-Object390 · 2026-07-28
- Nvidia puts an open vision-language-action model on Hugging Face — theteknosaur · 2026-07-28
- A Grok user asks to turn Hodor into Shrek in a classic AI meme prompt — heypearlai · 2026-07-28
- Fateward shows a cinematic RPG built with AI video — Kamacit · 2026-07-28
- Grok gets a Night King-to-Elsa crossover prompt in a meme-ready image edit — eyishazyer · 2026-07-28