Zero-WAM lets robots learn unseen tasks in-context from a single human video, hitting 46.95% zero-shot success

ZeYanjie · x · 2026-08-29

Core question

Can robots perform unseen tasks via in-context learning (ICL) like LLMs do? Zero-WAM, from Junwei Liang and colleagues, answers yes: it treats a human demonstration video as the prompt, specifying what to manipulate, how, and in what order.

Key components

Results

Zero-WAM reaches 46.95% zero-shot success on 7 unseen tasks, +29.5 points over Lingbot-VA. Paper, project page, and code are all public.

Related event: Zero-WAM: Robots Learn New Tasks Zero-Shot from Human Video Prompts(5 posts)→

Original post →

More from Embodied

Embodied channel →