CoRL 2026 paper GRA: synthetic robot videos should supervise geometry, not control
MikeShou1 · x · 2026-09-05
- 'Supervise What Survives' (arXiv 2606.24448) accepted to CoRL 2026, from Danze Chen et al. with Mike Zheng Shou
- Core insight — Asymmetric Preservation Principle: human-to-robot video generation preserves visible geometry (the 'where') while erasing control signals (the 'how'), so recovering pseudo-actions from synthetic pixels is a mismatched abstraction
- GRA extracts future 2D end-effector waypoints via pose estimation, retargeting, simulation and calibrated projection, feeding them through an auxiliary 2D head to supervise the VLA vision backbone; the action head trains only on real teleoperation demos
- Motivation: VLAs need massive video-action pairs but real teleoperation data is scarce; generated videos are a scalable, label-free alternative
More from Embodied
- Austin's locally made autonomous vehicle serving riders straight from the factory — yunta_tsai · 2026-09-05
- GeneralistAI collects 500,000+ hours of real robot data, launches onchain data bounties via Robinhood — broodsugar · 2026-09-05
- Lab's Three Bambu Printers Now Run Flat-Out As DNA Barcoding Scales to Thousands of Samples — generativist · 2026-09-05
- Tesla Cybercab Robotaxi Officially Launches at Austin Event — xfour · 2026-09-05
- Record 1,000+ Chinese firms at IFA2026 as humanoid robots take center stage — 智东西 · 2026-09-05
- Robotics data collection is now the most saturated startup idea, founder warns — arian_ghashghai · 2026-09-05