Berkeley's Do as I Do turns everyday human videos into dexterous robot hand training data
micoolcho · x · 2026-09-24
A UC Berkeley team (Paliwal, Etukuru, Liang, advised by Pieter Abbeel, Jitendra Malik, and Mahi Shafiullah) introduces Do as I Do, an algorithm that converts abundant monocular RGB human videos into robot-complete manipulation data.
Key points:
- Reconstructs hand-object interactions from egocentric and exocentric in-the-wild video sources
- Retargets these estimates into executable action sequences for multi-fingered dexterous robotic hands, bridging the human-to-robot embodiment gap
- Outperforms prior state of the art in hand-object interaction estimation and trajectory extraction, on ground-truth datasets and web-collected clips
- Demonstrated 10 real-world motions across varied object geometries and grasps
Paper and code are open-sourced, and the team proposes an efficacy playbook for practitioners collecting human manipulation data. A full RoboPapers podcast episode with the authors is dropping soon.
More from Embodied
- Daily AI brief: Meta glasses VR, GPT-6 voice plugins, Gemini TTS, Opus 5.5 — testingcatalog · 2026-09-24
- Meta Mocked for 'Inventing the Tamagotchi' as Users Say They Want One — jonstephens85 · 2026-09-24
- Robots fold laundry and write code, but camera calibration still takes hours — RexDouglass · 2026-09-24
- 36B Agent Model Runs Fully Local on Snapdragon X2 Elite With Just 32GB RAM — Kyrannio · 2026-09-24
- EgoSuite and continuous learning for robotics shown at Humanoids Summit Seoul, next at IROS — jonstephens85 · 2026-09-24
- CUHK's PackLab: an MLLM framework for closed-loop robotic bin packing beats RL — CUHK-CSE · 2026-09-24