Berkeley's HuGo uses LLMs to write humanoid policy code — no demos or reward design, 93.3% zero-shot on hardware
berkeley_ai · x · 2026-09-29
Researchers at UC Berkeley's ICON Lab with Google DeepMind collaborators present HuGo, a method where an LLM writes closed-loop high-level policy code for humanoid loco-manipulation tasks. Paired with a frozen low-level whole-body controller, it requires no teleop data, reward design, human motion data, or motion retargeting.
- Self-improving: the LLM analyzes its own simulation rollouts to refine policy code, and can also improve it directly on hardware
- Zero-shot transfer: policies generated in simulation run directly on hardware — 93.3% success on button pressing, 66.7% on box lifting
- From just a task description plus observation/command specs, the LLM decides where to stand, where to reach, and when to transition stages; code is coming soon
- Authors say the idea kept surprising them, and this predates Astra
More from Embodied
- SeeQ robot project retrospective: 'things just worked' when the ML was done right — aviral_kumar2 · 2026-09-29
- CMU's SeeQ: generalist robot value function on a VLM nearly doubles real-world task success — aviral_kumar2 · 2026-09-29
- Tech Workers Let ChatGPT, Claude and Grok Drive a Real Toyota Corolla; Only One Model Finished — 404 Media · 2026-09-29
- Forlinx's 20-TOPS M.2 AI Accelerator Supports PCIe Cascading for Local LLM Inference — DeliciousBelt9520 · 2026-09-29
- MolmoAct 2 tops robot board at 46% success with just 300 demos — DJiafei · 2026-09-29
- Rumor: OpenAI Pays $380k-$460k for Control Engineers to Bake Robotics into GPT-6 — xeophon · 2026-09-29