Vision-OPD uses 6.2K synthetic samples to boost fine-grained vision understanding

小红书技术REDtech · wechat · 2026-07-28

Xiaohongshu’s dots team introduces Vision-OPD, a vision-language post-training method designed to close the “region-to-global gap” in fine-grained visual understanding.

The paper argues that multimodal models often answer correctly when shown a cropped, zoomed-in region but fail on the full image. Vision-OPD turns that advantage into training signal by using an online self-distillation setup: a cropped-image teacher and a full-image student share the same model, while token-level divergence is minimized on the student’s own sampled trajectories.

Key results:

Original post →

More from Multimodal

Multimodal channel →