RLHND turns video diffusion models into physically grounded hand trackers for robot learning

RLWRLD · hf · 2026-10-08

RLHND repurposes a pre-trained Cosmos 3 video diffusion backbone into a deterministic clip-level feature extractor that jointly estimates hand pose and tactile information (contact and force) from monocular egocentric videos.

Key points:

Code will be open-sourced.

Original post →

More from Embodied

Embodied channel →