LeRF: Learning Reference Coordinate Frames for Perspective Taking Reasoning
Bang Xiao, Wenqi Jia, Ozgur Kara, Tiancheng Shen, Yibo Yang, Bolin Lai, Junho Kim, James Matthew Rehg
cs.CV
2026-09-29
LeRF trains Qwen3.5 VLMs to predict and render reference coordinate frames onto images before answering, +11 pp on hypothetical viewpoints over Qwen3.5-9B, ahead of GPT-5.6.
Perspective taking means answering spatial questions from a viewpoint other than the camera: another person's, or an imagined observer's. It is a base capability for navigation and human-robot collaboration, and VLMs fail at it in a consistent way, answering from the camera view no matter whose perspective the question asks about. GPT-5.6-Terra scores 81.37 on the camera-view (Ego) subset of OmniSpatial-PT, but only 55.85 on entity-centered (Allo) and 53.01 on hypothetical (Hypo) viewpoints.
Two existing fixes each carry a cost. SpatialReasoner estimates explicit 3D coordinates and derives directions geometrically, so 3D perception errors flow straight into reasoning. APC rebuilds a coarse 3D scene with external vision models and transforms it into the target viewpoint, adding multi-stage inference overhead and assuming every question needs a transform. The paper starts from a cheap observation: overlaying a projected reference frame from the off-the-shelf orientation model OrientAnything-V2 already fixes many failures. If drawing a frame helps, can the VLM learn to draw its own? That question produces LeRF, from UIUC, Amazon AGI and collaborators, with code released.
LeRF is built on Qwen3.5-4B/9B and runs a two-turn tool-call loop.
Training is two-stage. SFT derives frame labels from object and human pose datasets (ImageNet3D, Omni6DPose-SOPE, BEDLAM), with GPT-5.6-Terra generating referring descriptions and filtering ambiguous samples; about 20% of the data teaches when to skip the tool. LoRA rank 8, vision encoder and projector frozen. RL then applies GRPO on perspective-taking VQA (MultihopSpatial-train, 6.79K queries; SpatialReasoner-RL, 1.2K pairs) with a binary reward on final-answer correctness only. Frame-prediction tokens are excluded from the policy gradient so RL does not erode what SFT taught. All training runs on 4 RTX PRO 6000 Blackwell GPUs.
| Method | OmniSpatial Allo | OmniSpatial Hypo | 3DSRBench Ori | V-Spatial P-Rel |
| Qwen3.5-9B | 47.13 | 44.58 | 48.17 | 65.51 |
| GPT-5.6-Terra | 55.85 | 53.01 | 63.32 | 77.20 |
| LeRF-9B | 54.04 | 55.66 | 53.76 | 74.23 |
LeRF-9B tops all open-source methods on the five viewpoint-change subsets (Allo, Hypo, 3DSRBench Orientation, both ViewSpatial-Bench tasks) and places second on 3DSRBench Multi-Object. On Hypo (55.66) and ViewSpatial-Bench P-Obj (61.91) it passes every proprietary model evaluated, GPT-5.6 and Claude Sonnet 5 included. Against its backbone the gains are +11.08 pp on Hypo, +6.91 on Allo, +8.72 on P-Rel. The trade-off: Ego drops from 80.20 to 74.31, which the paper acknowledges.
Frame estimation itself gets much stronger. On the EMDB human subset, LeRF-9B lands all three projected axes within 15° on 43.12% of views, 23.77 pp above the specialist orientation model OrientAnything-V2 (19.35%), with 99.82% origin localization; the raw Qwen3.5-9B backbone manages 1.11%. On OmniNOCS-Objectron objects it nearly matches OrientAnything-V2 at All@15° (49.58 vs 50.42) and trails at All@30° (72.54 vs 83.23).
Two analyses support the mechanism. Rotating every rendered frame 90° counterclockwise at test time drops overall accuracy from 57.97% to 54.05%, with the largest fall on Hypo (7.95 pp): the model genuinely reads the drawn axes. Tool use is selective, 2.94% on egocentric questions versus above 84% on the other two types.
Ablations: the SFT-only checkpoint falls to 40.86% overall, below the base model's 52.76%, so SFT alone damages general reasoning. But it initializes RL well: SFT then RL reaches 57.97% versus 53.76% for RL straight from the base. Rendering matters most for hypothetical viewpoints, 55.66 with rendering in both training and eval versus 49.40 with neither.
LeRF recasts an implicit mental rotation as grounded visual perception, and at inference it needs only a deterministic renderer: no external perception models, no 3D reconstruction. A 9B backbone, LoRA, and four GPUs put reproduction within reach of most labs. For embodied AI and robotics teams the pattern transfers directly: teach the model to build a lightweight structured intermediate representation instead of stacking external modules. The training recipe, SFT for tool use plus RL with an answer-only reward and tool-call tokens masked from the gradient, applies to tool-use training well beyond spatial reasoning.
Keep expectations calibrated: gains over APC and SpatialReasoner are clear, but GPT-5.6-Terra still leads on Allo, both 3DSRBench categories, and P-Rel.
The authors flag occlusion, visual ambiguity, and uncertain entity orientation as error sources that propagate into reasoning, and the method handles static images only; dynamic scenes with moving reference frames are untested. Further concerns from the results: Ego accuracy drops 5.89 pp, a net loss for camera-centric applications; rendering gains concentrate on Hypo (about 6 pp) while the overall difference is 0.79 pp, so the claim that rendering is indispensable holds mainly for hypothetical viewpoints; and the RL stage ran 50 steps, leaving longer-training behavior unmeasured.