Geometry-Privileged Distillation lifts VLM spatial reasoning while keeping RGB-only deployment
zju · hf · 2026-10-09
ZJU's GPD uses 3D evidence (depth, semantic, BEV cues rendered as text) as privilege in on-policy self-distillation, augmenting GRPO with privileged KL applied only to incorrect trajectories—deployed models stay RGB-only. The 4B backbone reaches 57.1 on VSI-Bench and 37.6 average across four spatial benchmarks, beating GRPO and answer-privileged OPSD.
More from Models
- Open-Source NSFW Classifier Blue-Eye Hits 88.9%, Beats AWS and Google Vision — Mundane_Toe_8074 · 2026-10-09
- Every model looked bad in my eval — the bug was my answer key, not the models — jgarg27 · 2026-10-09
- Ai2: the next Olmo model is already training, fully open-model commitment unchanged — sewon__min · 2026-10-09
- Whistle: an open 16.9 MB speech-to-text model that runs on CPU with 11 ms first token in 7 languages — solyarisoftware · 2026-10-09
- Nous Research's typo 'Hemres Agnet' looks like a Hermes Agent teaser — Teknium · 2026-10-09
- Study: Outdated Gemini 2.5 Advice Rated on Par With Doctors in Urgent Care, No Safety Issues — emollick · 2026-10-09