Zero-WAM lets robots learn unseen tasks in-context from a single human video, hitting 46.95% zero-shot success
ZeYanjie · x · 2026-08-29
Core question
Can robots perform unseen tasks via in-context learning (ICL) like LLMs do? Zero-WAM, from Junwei Liang and colleagues, answers yes: it treats a human demonstration video as the prompt, specifying what to manipulate, how, and in what order.
Key components
- HumanGen pipeline: automatically converts task-sampled robot trajectories into semantically matched human demonstrations, yielding 74.2K human-robot pairs across 8.6K tasks to address data scarcity.
- Model: a causal video-action model that executes unseen tasks by following in-context human video guidance.
- IFP objective (in-context future chunk prediction): suppresses shortcuts learned from seen tasks and forces the policy to draw task information from the video prompt.
Results
Zero-WAM reaches 46.95% zero-shot success on 7 unseen tasks, +29.5 points over Lingbot-VA. Paper, project page, and code are all public.
Related event: Zero-WAM: Robots Learn New Tasks Zero-Shot from Human Video Prompts(5 posts)→
More from Embodied
- REAL: an embodied agent that explores real rooms and asks clarifying questions, accepted at ECCV 2026 — jiqizhixin · 2026-09-23
- Tipping a locked-standing humanoid upright is how every robot on the market gets up, says operator — carlosdponx · 2026-09-23
- Why a humanoid robot auto-damps in locked standing: stiff joints shatter on falls — carlosdponx · 2026-09-23
- An AI-powered BD company to help robotics firms find high-value use cases — gan_chuang · 2026-09-23
- iPhone 18 Pro's variable aperture fixes the 3D spatial video problem that plagued 17 Pro — Scobleizer · 2026-09-23
- SPECS Glasses Demo Core Boom, a Spatial Shooter Built for AI Wearables — TinfoilTricorn · 2026-09-23