HUG uses 1M egocentric frames to train zero-shot robot grasping
chris_j_paxton · x · 2026-07-25
RoboPapers highlights HUG (Human Universal Grasping), a method for learning robot grasping from egocentric human video alone.
Key points:
- The team collects 1M frames / 27.8 hours of egocentric human grasping video.
- They train a flow-matching model to predict human hand pose.
- Those predicted hand poses are then retargeted to robot hands.
- The approach reportedly brings large gains on a range of zero-shot robot grasping tasks in everyday scenes.
The episode features @kevinywu, @irmakkguzey, and @DandanShan discussing the work and why human video may be a scalable substitute for scarce robot data.
More from Embodied
- Eyecandy Robotics plans a $200 palm-sized tabletop robot for 2027 — k7agar · 2026-07-25
- An $8 ESP32-S3 now runs a fully offline, real-time object recognition camera — Yamapama · 2026-07-25
- SpaceX says Starship V3 is improving reusable orbital heat-shield tiles flight by flight — XFreeze · 2026-07-25
- Reddit argues robotics is held back more by AI than by hardware — ErmingSoHard · 2026-07-25
- Prototype AI drone patrols a yard, spots dog poop, and scoops it up — brianrkelly · 2026-07-25
- Hardware hacker 3D-prints a chess set with an AI mechanical engineer — dee_hw · 2026-07-25