Cornell's UMA robot model learns new tasks from 25 human videos with frozen weights
qinzytech · x · 2026-10-06
Cornell's Unified Motion-Action (UMA) model shows a robot can learn a new task from just 25 human video demonstrations with model weights frozen — adaptation only updates a small set of task tokens representing what to do.
- Approach: UMA bridges human videos and robot control via 3D object motion (how points on objects move over time). It learns motion from videos without robot action labels, and learns the motion-action link from robot data; pretraining mixes human videos, real robot demos, and simulation.
- Results: in video-only adaptation, UMA beat UVA by 25 percentage points on each task (utensil insertion, sweeping, folding jeans; 20 trials each), with human demos from the evaluation scene and viewpoint.
- Takeaway: new tasks can be taught via human demos without collecting robot action labels, though deploying on a new embodiment still needs compatible robot data; cross-scene/viewpoint transfer remains an open test.
More from Embodied
- Watch, Infer, Coordinate: robots infer a partner's physical limits from watching teamwork, then coordinate zero-shot — mangahomanga · 2026-10-06
- Embodied Analysis launches unified eval platform for physical agents and world models — qinzytech · 2026-10-06
- T²Mem lifts robot policy π₀.₅ success rate from 17.93% to 56.83% on memory tasks — yining_hong · 2026-10-06
- Atomic-layer-deposited titanium-platinum actuators that "walk": a talk with only ~2k views — Promptmethus · 2026-10-06
- Manda Robotics benchmarks Newton vs MuJoCo vs Genesis on deformable objects — facontidavide · 2026-10-06
- Morgan Stanley: 1 billion humanoids could be in use within 25 years — Rewkang · 2026-10-06