LeRF teaches VLMs explicit reference frames to crack perspective-taking reasoning

UIUC-CS · hf · 2026-09-30

Researchers propose LeRF (Learning Reference Coordinate Frames), tackling the weakness of VLMs in perspective taking, where models default to the camera viewpoint instead of reasoning from another entity's perspective.

Original post →

More from Research

Research channel →