TANGO: whole-body VLA lets humanoid robots zero-shot navigate cluttered spaces on Unitree G1
chris_j_paxton · x · 2026-09-09
A CoRL 2026 paper introduces TANGO, a whole-body vision-language-action framework for humanoid navigation: from a language instruction and egocentric RGB, it directly predicts 29-DoF joint actions coordinating arms, torso, and gait to traverse cluttered 3D indoor spaces.
- Navigation is framed as a whole-body problem rather than a 2D abstraction.
- Trained entirely in simulation via a Plan-Edit-Track pipeline (global path planning, kinematic whole-body motion generation, obstacle-aware editing, RL tracking).
- Triple-system architecture: vision-language backbone, action expert, and tracking, producing executable whole-body action chunks.
- Zero-shot transfer to a real Unitree G1 with no real-world data. Commenter Chris Paxton notes perception-aware whole-body control is crucial yet still rare.
Related event: TANGO: Whole-Body VLA Enables Humanoid Robots to Navigate Cluttered Spaces(2 posts)→
More from Embodied
- Musk: Cybercab certified at 165 Wh/mi, the most efficient production EV ever — elonmusk · 2026-09-11
- MKBHD goes hands-on with the first folding iPhone; $2,000 48MP selfie cam mocked — alexmacgregor__ · 2026-09-11
- NTU spin-off Ropedia launches HOMIE Gen 2 wearable system to train robots from human experience — liuziwei7 · 2026-09-11
- M.2-Oculink eGPU Link Silently Downgrades to PCIe Gen1 — Here's How to Check — El_90 · 2026-09-11
- Replit Agent in a robot builds and publishes websites autonomously via MCP — amasad · 2026-09-11
- TARS Robotics unveils embodied foundation model AWE: 15+ tasks, one model, zero retraining — heyshrutimishra · 2026-09-11