Stanford's HomeBody Lets GPT Astra Directly Orchestrate a Humanoid Robot — No VLA Training Needed

ZeYanjie · x · 2026-09-26

Researchers from Stanford and Caltech unveiled HomeBody, a system that tests whether humanoid robots still need a learned VLA layer between reasoning and motion.

The idea: The standard stack is System 2 VLM (reasoning) → System 1 learned VLA → System 0 motion controller. HomeBody removes the learned VLA entirely, letting a frontier VLM (GPT Astra in their setup) directly call a library of composable motor skills.

Results: On a Unitree G1 in a previously unseen kitchen, the robot completed long-horizon tasks — tidying across the room and retrieving a remembered object from an ambiguous request — with no environment-specific training data or extra policy learning. The VLM gets persistent spatial memory and revises decisions from skill execution feedback.

Paper and code are open-sourced, suggesting "frontier VLM + skill library" may replace end-to-end learned action pipelines.

Original post →

More from Embodied

Embodied channel →