Where-OPD: synthetic-scene spatial self-distillation boosts MLLM perception by 3.23 points
valeocorg · hf · 2026-10-02
The Where-OPD paper introduces spatially guided on-policy self-distillation for multimodal LLMs. Instead of human-annotated grounding data or external teachers, the EMA teacher receives textual spatial guidance on query-relevant visual elements, trained on procedurally generated synthetic scenes with free object identities and coordinates.
Key points:
- Teacher locates and integrates evidence across image regions; the student learns to reproduce this from image and question alone
- Post-training uses only synthetic scenes, yet gains transfer to real-world perception benchmarks
- +3.23 points average across CVBench, V, ZoomBench, BLINK, HR-Bench, and MME-RealWorld
Project page: github.com/sirkosophia/Where-OPD
More from Multimodal
- Full music video generated locally with ComfyUI and LTX 2.5 on 16GB VRAM — sokmech · 2026-10-02
- Eleven v4 character animation demo shown off in new video — buraktuyan · 2026-10-02
- Claude directs its own EDM music video 'The Good Ending' in Fable 5.1 demo — cheetoskull · 2026-10-02
- All-in-one Krea 2 Turbo ComfyUI workflow uses 2/4-step distilled LoRAs for speed — TimeTruth2490 · 2026-10-02
- No one on the team knew Blender — Claude ran the whole 3D film workflow — gen_ericai · 2026-10-02
- Meshy teases Meshy Edit: tweak specific parts of 3D models with a text prompt — rms80 · 2026-10-02