Where-OPD: synthetic-scene spatial self-distillation boosts MLLM perception by 3.23 points

valeocorg · hf · 2026-10-02

The Where-OPD paper introduces spatially guided on-policy self-distillation for multimodal LLMs. Instead of human-annotated grounding data or external teachers, the EMA teacher receives textual spatial guidance on query-relevant visual elements, trained on procedurally generated synthetic scenes with free object identities and coordinates.

Key points:

Project page: github.com/sirkosophia/Where-OPD

Original post →

More from Multimodal

Multimodal channel →