HUG learns robot grasping from 1 million human handshapes and beats baselines by 34%
LerrelPinto · x · 2026-07-29
The team behind HUG (Human Universal Grasping) flew their robot to Seattle for the Meta ARIA Summit to show that the system works in the real world.
- HUG is a flow-matching model that predicts diverse human-style grasps from a single RGB-D image.
- The training data, 1M-HUGs, contains 1 million egocentric grasp frames, 27.8 hours of video, 6,707 object instances, and 41 buildings.
- The model retargets grasps to different robot hands, enabling zero-shot grasping in everyday scenes.
- The authors also release a new benchmark, HUG-Bench, with 90 unseen objects, plus code, weights, dataset, and an interactive demo.
- On the real-world 30-object test set, HUG beats prior baselines by 23% and 34% on the challenging object set.
More from Embodied
- Directory of US Quick-Turn Hardware Suppliers Amid FCC Robot Bans — mattfreed · 2026-07-29
- Khora proposes a multi-agent world model that scales linearly instead of quadratically — siyuanhuang95 · 2026-07-29
- Nature Communications: Spatial Network Principles Underlying Neural Locomotion — plopesresearch · 2026-07-29
- China shows a wall-climbing embodied robot as BYD, Honor and Arm push new hardware — 创业邦 · 2026-07-29
- Seattle demo shows a rideable robot and autonomy for scooters and bikes — Scobleizer · 2026-07-29
- A one-camera robot folds a shirt zero-shot with just an RGB-D sensor — Scobleizer · 2026-07-29