HO-Cap: Multi-View RGB-D Hand-Object Interaction Dataset Without Motion Capture

YuXiang_IRVL · x · 2026-08-16

HO-Cap is a dataset for 3D reconstruction and pose tracking of hand-object interaction, released by UT Dallas and NVIDIA. The system uses multiple RGB-D cameras and a HoloLens headset for data collection, avoiding expensive 3D scanners or mocap systems. A semi-automatic annotation method significantly reduces annotation time. The dataset includes videos of humans interacting with objects for tasks like pick-and-place, handovers, and affordance usage, serving as human demonstrations for embodied AI and robot manipulation.

Original post →

More from Embodied

Embodied channel →