LeRF teaches VLMs explicit reference frames to crack perspective-taking reasoning
UIUC-CS · hf · 2026-09-30
Researchers propose LeRF (Learning Reference Coordinate Frames), tackling the weakness of VLMs in perspective taking, where models default to the camera viewpoint instead of reasoning from another entity's perspective.
- Method: Given an image and query, the model decides whether a coordinate frame is needed; if so, it grounds the reference entity and predicts the frame's origin and entity-centered frame, then a lightweight renderer overlays it onto the image — no external perception models or explicit 3D reconstruction required.
- Training: Supervised fine-tuning for selective tool invocation and frame prediction, followed by reinforcement learning on spatial VQA pairs.
- Results: Consistent gains over the backbone across perspective-taking benchmarks, plus improved reference-frame grounding and orientation estimation versus open-source methods.
More from Research
- AutoRef open-sourced: harness optimization for agentic multi-reference image generation — NunyaBuzor · 2026-09-30
- VLANeXt Family: 500+ controlled experiments distill 12 practical recipes for VLA models — ccloy · 2026-09-30
- RL with Confidence Margin: COLM 2026 paper makes step-by-step confidence track reasoning correctness — EliasEskin · 2026-09-30
- Alibaba DAMO unveils WorldAttention for efficient interactive video world models — Alibaba-DAMO-Academy · 2026-09-30
- ActFirst-OPD trains multi-turn agents up to 4.9x faster by acting before reasoning — SouthernUniversityofScienceandTechnology · 2026-09-30
- NUS's MaLiang-Harness exposes the Program-to-Visual gap in code-driven image/video generation — NationalUniversityofSingapore · 2026-09-30