TANGO: Sim-Only Trained Whole-Body VLA Gives Humanoids Zero-Shot Navigation on Unitree G1
chris_j_paxton · x · 2026-09-24
CoRL 2026 paper TANGO introduces a whole-body vision-language navigation framework that predicts 29-DoF joint-space actions directly from egocentric RGB, letting humanoids navigate cluttered 3D spaces with coordinated arms, torso, and gait.
- Treats navigation as a whole-body problem requiring continuous geometry-aware adaptation, not a 2D abstraction
- Plan-Edit-Track pipeline synthesizes diverse collision-free traversal data in simulation via global path planning, kinematic motion generation, obstacle-aware editing, and RL-based tracking
- Triple-system model (vision-language backbone + action expert) outputs executable whole-body action chunks
- Trained entirely in simulation, transfers zero-shot to a real Unitree G1 with no real-world training data
The authors note large-scale sim for training policies is really starting to work.
More from Embodied
- Commercial autonomous trucking is officially here — SuB8u · 2026-09-24
- Smart glasses are already causing havoc in India — and a crackdown is unlikely — krishnan · 2026-09-24
- RealSense VP on AgenticROS: Letting AI Agents Directly Control Physical Robots — chrismatthieu · 2026-09-24
- Scoble: the future of XR is glasses, not scuba-mask headsets — Scobleizer · 2026-09-24
- Snapdragon X2 mid-cycle refresh: new Surface, Chromebooks, and Linux support — ryanshrout · 2026-09-24
- AREX launches AR scuba mask with dive data in your line of sight, $899 and up — Scobleizer · 2026-09-24