Robot-centric pointmaps improve VLA policies
kaist-ai · hf · 2026-07-20
KAIST AI proposes robot-centric pointmaps for vision-language-action (VLA) models.
What problem it solves
VLAs usually observe scenes in the camera frame, but robot actions are defined in the robot’s own 3D frame. That mismatch becomes harder when training data comes from many different camera viewpoints.
What the method does
- Uses images whose pixels store 3D coordinates in the robot frame
- Preserves the dense H × W grid expected by pretrained 2D VLAs
- Integrates into existing VLAs with minimal architectural changes
Results
- On RoboCasa, pointmaps improve both pi0.5 and SmolVLA
- They outperform camera-viewpoint and 3D-aware baselines
- In real-robot tests, the advantage grows when the camera is moved to an unseen placement
More from Embodied
- NVIDIA brings its Cosmos 3 Edge world model to Jetson for on-device robot control — liu_mingyu · 2026-07-21
- A set of agent skills for CAD, robotics, and hardware design — earthtojake · 2026-07-21
- DIY wooden box packs 6 Intel Arc Pro B70 cards with a FreeCAD model — nick_ziv · 2026-07-21
- Snake-like robot moves on fully passive wheels and winding motion — ___Mufasaa · 2026-07-21
- Creator buys a Reachy robot and asks what to build first — dee_hw · 2026-07-21
- MW team shows Gen1 of MW-bot, a semi-humanoid home robot built for pantry storage and ceiling rails — CyberRobooo · 2026-07-21