Human Universal Grasping
Kevin Yuanbo Wu, Tianxing Zhou, Isaac Tu, Billy Yan, Irmak Guzey, David Fouhey, Dandan Shan, Lerrel Pinto
cs.RO, cs.AI, cs.CV, cs.LG
2026-06-16
HUG trains flow matching on 1M human grasps, predicts MANO, and retargets to robots. Tabletop success on 30 unseen objects is 66.7%, 23 points above Dex1B.
People pick up thousands of objects a day. Multi-fingered robots do not. Data is the bottleneck. Simulated force-closure search and RL sampling hit the sim-to-real gap and usually need retraining per hand. Teleoperation yields real grasps on the target hardware but is slow and cannot cover the open world.
HUG's claim is that the natural source of grasp data is humans. Aria Gen 2 glasses stream calibrated RGB-D and hand landmarks; anthropomorphic hands plus learned retargeting shrink the morphology gap. That unlocks a pipeline of collect, learn, retarget, with no robot data in training.
The dataset is 1M-HUGs. A wearer circles a static object for 15 to 30 seconds with no hand in view, then reaches in with the right hand. Camera poses copy that one grasp back onto the earlier empty frames, so one physical grasp becomes hundreds of (object-only image, grasp) pairs. After filtering there are 1M RGB and 1M grayscale frames, about 1.5K unique objects, 41 buildings, and 6,707 recordings. Twenty-one landmarks are fit to MANO with shape frozen, so pose is comparable across people, loadable in simulation, and retargetable.
The model is point-conditioned flow matching. A click on the RGB-D image yields a 99-d grasp: wrist translation, 6D wrist rotation, and 6D rotations for 15 finger joints. Frozen DINOv2 encodes RGB. PointNeXt encodes a 0.3 m crop of the metric point cloud. Point painting pastes DINO features onto cloud centroids. A 4-layer fusion transformer feeds a 6-layer DiT that predicts velocity on three token groups: translation, wrist, fingers. The loss is velocity MSE plus a 3D fingertip term weighted 20, stronger near clean samples. Training takes about 10 hours on two RTX 5090s.
At deployment the MANO grasp is retargeted to Ability, WUJI, and similar hands and executed open-loop as pre-grasp, close, lift. No per-hand training.
HUG-Bench has 90 unseen objects across five geometries and three sizes, including objects about 1 cm tall and large awkward ones. Simulation test success is 73.0% against a human-replay oracle of 94.0%. Dropping the 3D loss falls to 32.7%. RGB-only is 29.7%; point-cloud-only is 70.7%. Scaling data from 25K to 1M frames lifts test success from 33% to 73% with no plateau.
Real-world, 30 objects, 10 trials each:
| Method | Tabletop success | ≥1 success |
| CAP (parallel jaw) | 32.7% | 20/30 |
| Dex1B | 43.7% | 27/30 |
| HUG | 66.7% | 28/30 |
| HUG in-the-wild (new arm and camera) | 62.0% | 29/30 |
Storage bin 10/10, picnic basket 9/10, spray bottle 9/10. Football and wipe dispenser 0/10: the Ability hand cannot wrap them. Most failures hit the object or table while closing.
Dexterous grasping does not have to wait on robot teleoperation. Human grasp distributions are more executable than the set of all physically valid grasps, and retargeting attaches one predictor to many hands. Code, data, the benchmark, and a demo are public. For household robots this chain is closer to usable than another billion simulated grasps.
Right hand only, frozen MANO shape, no left or bimanual grasps. Execution is open-loop with no visual feedback in contact. Occluded hands make Aria tracking too loose or too tight. Input is 224×224, so tiny objects and large distant ones suffer. Evaluation is indoor. A rejected CoRL rumor is irrelevant here; the paper already plots the failure modes. Motion planning and force-aware closing are next steps, not solved claims.