Zero-WAM: robots learn unseen tasks from a single human demo video, no fine-tuning

deepakpathak · x · 2026-08-28

Researchers from HKUST (GZ) and Robbyant present Zero-WAM, which treats a human demonstration video as an in-context prompt: the model watches how the scene should evolve and generates robot actions without task-specific fine-tuning.

Motivation: language says "put this there," but video also shows path, timing, contact order, and what "there" actually looks like.

Related event: Zero-WAM: Robots Learn New Tasks from Human Videos Without Fine-tuning(2 posts)→

Original post →

More from Embodied

Embodied channel →