Geometry-Privileged Distillation lifts VLM spatial reasoning while keeping RGB-only deployment

zju · hf · 2026-10-09

ZJU's GPD uses 3D evidence (depth, semantic, BEV cues rendered as text) as privilege in on-policy self-distillation, augmenting GRPO with privileged KL applied only to incorrect trajectories—deployed models stay RGB-only. The 4B backbone reaches 57.1 on VSI-Bench and 37.6 average across four spatial benchmarks, beating GRPO and answer-privileged OPSD.

Original post →

More from Models

Models channel →