Stanford's HomeBody Lets GPT Astra Directly Orchestrate a Humanoid Robot — No VLA Training Needed
ZeYanjie · x · 2026-09-26
Researchers from Stanford and Caltech unveiled HomeBody, a system that tests whether humanoid robots still need a learned VLA layer between reasoning and motion.
The idea: The standard stack is System 2 VLM (reasoning) → System 1 learned VLA → System 0 motion controller. HomeBody removes the learned VLA entirely, letting a frontier VLM (GPT Astra in their setup) directly call a library of composable motor skills.
Results: On a Unitree G1 in a previously unseen kitchen, the robot completed long-horizon tasks — tidying across the room and retrieving a remembered object from an ambiguous request — with no environment-specific training data or extra policy learning. The VLM gets persistent spatial memory and revises decisions from skill execution feedback.
Paper and code are open-sourced, suggesting "frontier VLM + skill library" may replace end-to-end learned action pipelines.
More from Embodied
- German startup ecoro pilots fully autonomous inter-building cargo system in Japan — 4310sy · 2026-09-26
- Reachy Pilot links SPECS smart glasses to a Reachy Mini robot for remote guarding — Scobleizer · 2026-09-26
- ProxyPose lands at NeurIPS: track 6-DoF pose from a single query pixel — CSProfKGD · 2026-09-26
- Using ChatGPT Voice to query pesticide records: it reads and updates Airtable hands-free — athyuttamre · 2026-09-26
- Vibe-coding hardware: build an ESP32 e-ink dashboard with Codex and Cloudflare — paw_lean · 2026-09-26
- Niantic opens Places Library: 100 real-world 3D scenes for embodied AI training — Scobleizer · 2026-09-26