RLHND turns video diffusion models into physically grounded hand trackers for robot learning
RLWRLD · hf · 2026-10-08
RLHND repurposes a pre-trained Cosmos 3 video diffusion backbone into a deterministic clip-level feature extractor that jointly estimates hand pose and tactile information (contact and force) from monocular egocentric videos.
Key points:
- Motivation: existing hand trackers regress pose from cropped frames with weak priors on hand motion and object interaction, and lack physical cues needed for robot policy training from human videos.
- Method: clean-latent conditioning carries video priors into tracking; a pose stream predicts anatomically plausible joints with optional shape conditioning; a separate tactile expert stream (trained with pose frozen) predicts dense contact and force; LBS-based feature spreading avoids costly per-vertex attention.
- Results: state-of-the-art on multiple benchmarks for pose, contact, and force estimation, plus retargeting and real-world robot experiments.
Code will be open-sourced.
More from Embodied
- NEEDLEwork stitches suboptimal robot demos into better training data — SongShuran · 2026-10-08
- Stanford's MobileVISTA fixes mobile manipulation's pose brittleness with generative data augmentation — jiajunwu_cs · 2026-10-08
- Researchers compile a larval fish brain into a simulated circuit running 26x faster than spiking models — mtizard · 2026-10-08
- US Treasury issues first outbound investment fine: $200k over $92k Chinese robotics AI deal — pstAsiatech · 2026-10-08
- Palmer Luckey: I Believe in Humanoid Robots, But Won't Build Them Myself — PeterDiamandis · 2026-10-08
- Adding Motion Sense to VLAs With Frozen SAM 2 Lets Robots Catch Moving Bottles — nikhilaravi · 2026-10-08