Robot-centric pointmaps improve VLA policies
kaist-ai · hf · 2026-07-20
KAIST AI proposes robot-centric pointmaps for vision-language-action (VLA) models.
What problem it solves
VLAs usually observe scenes in the camera frame, but robot actions are defined in the robot’s own 3D frame. That mismatch becomes harder when training data comes from many different camera viewpoints.
What the method does
- Uses images whose pixels store 3D coordinates in the robot frame
- Preserves the dense H × W grid expected by pretrained 2D VLAs
- Integrates into existing VLAs with minimal architectural changes
Results
- On RoboCasa, pointmaps improve both pi0.5 and SmolVLA
- They outperform camera-viewpoint and 3D-aware baselines
- In real-robot tests, the advantage grows when the camera is moved to an unseen placement
More from Embodied
- Johns Hopkins Launches Full-Stack Hands-on Robot Learning Class with SO-101 Arm Kits — _krishna_murthy · 2026-09-11
- SyncWorld: In-Context Robot World Model Simulates Unseen Views and Embodiments Zero-Shot — ChongZzZhang · 2026-09-11
- A 3D Pose Dataset for Dogs Released — ducha_aiki · 2026-09-11
- Swaayatt demos autonomous driving at 52 km/h on mountain roads, self-recovers after skid — sanjeevs_iitr · 2026-09-11
- AUAR's MicroFactory brings a deployable robotic wood-panel factory to the construction site — lukas_m_ziegler · 2026-09-11
- Musk: Cybercab certified at 165 Wh/mi, the most efficient production EV ever — elonmusk · 2026-09-11