TANGO: whole-body VLA model navigates humanoid robots in cluttered spaces from sim data

Anqi Li · hf · 2026-09-09

TANGO is a vision-language-action framework that predicts whole-body joint actions for humanoid robots navigating cluttered indoor environments, trained entirely on simulated data.

Original post →

More from Embodied

Embodied channel →