TANGO: Sim-Only Trained Whole-Body VLA Gives Humanoids Zero-Shot Navigation on Unitree G1

chris_j_paxton · x · 2026-09-24

CoRL 2026 paper TANGO introduces a whole-body vision-language navigation framework that predicts 29-DoF joint-space actions directly from egocentric RGB, letting humanoids navigate cluttered 3D spaces with coordinated arms, torso, and gait.

The authors note large-scale sim for training policies is really starting to work.

Original post →

More from Embodied

Embodied channel →