Robot-centric pointmaps improve VLA policies

kaist-ai · hf · 2026-07-20

KAIST AI proposes robot-centric pointmaps for vision-language-action (VLA) models.

What problem it solves

VLAs usually observe scenes in the camera frame, but robot actions are defined in the robot’s own 3D frame. That mismatch becomes harder when training data comes from many different camera viewpoints.

What the method does

Results

Original post →

More from Embodied

Embodied channel →