HuGo Uses LLMs and VLMs to Generate Robot Policies — No Demos or Reward Shaping Needed
m_wulfmeier · x · 2026-10-02
Most robot learning effort goes into collecting demonstrations or hand-engineering task-specific rewards. HuGo takes a different route: LLMs and VLMs act as general-purpose tools for policy generation and verification, requiring only language to define new tasks.
Key points
- LLM/VLMs produce high-level commands executed by NVIDIA's SONIC, with low-level RL controllers
- After extensive sim2real and continued real-world learning, the authors report striking data efficiency
- The same approach works out of the box for real-world adaptation
The team, which experienced the pain firsthand working on robot soccer, says the work was led by seoyeonchoi827.
More from Embodied
- LeKiwi robot arm swaps OAK-D S2 for Orbbec Gemini 2 camera — kamathsblog · 2026-10-02
- Shield AI borrows Tesla's simulation playbook, bets Aechelon deal wins AI pilot training — BrettKrieger12 · 2026-10-02
- OpenAI gifts developers a 'Codex Micro' hardware device, first units shown off — OpenAIDevs · 2026-10-02
- Bullish analyst: Cybercab registrations climbing, TSLA undervalues the buildout — JOBhakdi · 2026-10-02
- Rhoda's robots unpack, scan and stack component reels for real customers — vincesitzmann · 2026-10-02
- Agility's Digit picks totes autonomously at IROS with ~80-90% success — jonstephens85 · 2026-10-02