Zero-WAM: robots learn unseen tasks from a single human demo video, no fine-tuning
deepakpathak · x · 2026-08-28
Researchers from HKUST (GZ) and Robbyant present Zero-WAM, which treats a human demonstration video as an in-context prompt: the model watches how the scene should evolve and generates robot actions without task-specific fine-tuning.
- To train at scale, they built HumanGen, 74.2k generated human-robot pairs across 8.6k tasks.
- On 7 unseen RoboTwin tasks it hits 47.0% success, 29.5 points above the strongest video-action baseline, with real-robot transfer to unseen multi-object, long-horizon and insertion tasks.
Motivation: language says "put this there," but video also shows path, timing, contact order, and what "there" actually looks like.
Related event: Zero-WAM: Robots Learn New Tasks from Human Videos Without Fine-tuning(2 posts)→
More from Embodied
- AI agents make sim-to-real training more accessible — chris_j_paxton · 2026-08-28
- 2-year-old names 6 robots, signaling a future of normalized human-robot coexistence — chris_j_paxton · 2026-08-28
- U-BOT body width reduced to 250mm, 11.5-hour print completed — _Stocko_ · 2026-08-28
- Microduck AR runs a real robot RL policy in your phone browser, no server — multimodalart · 2026-08-28
- Pixel 11 Pro Integrates Gemini Nano 4 for Local Troubleshooting and 120x Zoom — Resplendent-Sun · 2026-08-28
- Opinion: AI's Next Frontier Is the Physical World, Not the Chatbot — ingliguori · 2026-08-28