ViDiHand: Video Diffusion Models Prove Surprisingly Good at 4D Hand Motion Reconstruction
ccloy · x · 2026-10-06
NTU MMLab and collaborators introduce ViDiHand, the first method to leverage a pretrained video diffusion model for 4D two-hand motion reconstruction from egocentric video.
- Why it matters: Egocentric human video is becoming a key supervision source for robot dexterous manipulation (EgoDex, EgoVLA, etc.), but heavy occlusion breaks existing image- and temporal-based methods.
- Approach: Instead of using the diffusion model as a frozen feature extractor, they adapt it via a hand-overlay rendering task, then a dual-branch decoder reads articulated 3D pose and per-joint 2D localization at metric scale.
- Paper, project page and code are public, with a follow-up ACE-Ego-Hand on Hugging Face.
More from Embodied
- VC: almost none of real robot deployment data flows back into training — carrycooldude · 2026-10-06
- RealtimeWAM: One-Step Asynchronous World Action Model Delivers 25x Speedup with <1% Accuracy Loss — NanyangTechnologicalUniversity · 2026-10-06
- ACG-Bench Probes Dual-Arm VLA Generalization; AE-VLA Lifts Success from 3% to 21.5% — Zaibin Zhang · 2026-10-06
- ReSteer open-sourced: fixing VLA policies that ignore mid-execution instruction switches — siddkaramcheti · 2026-10-06
- DIY robotic backboard uses computer vision to make every beginner shot go in — TinfoilTricorn · 2026-10-06
- Xitac tactile sensor uses magnetotactic gel over 3D hall array to capture finger forces — seanmcdonaldxyz · 2026-10-06