Human video emerges as robotics pretraining substrate: effective experience = hours × information per hour

rohanpaul_ai · x · 2026-09-22

Maxinsights argues human egocentric video is now a serious pretraining substrate for robot manipulation, but works best when translated into the robot's own embodiment first. It proposes scaling Physical AI data on two coupled axes: Experience Scale (total hours recorded) and Experience Density (learnable physics per hour — object states, contact, forces, geometry, tool use).

Evidence: Dyna-2, pretrained on 2M+ hours of egocentric video, improves monotonically across four orders of magnitude of human experience and transfers across the embodiment gap. Two recordings of equal duration can differ an order of magnitude in learning value — repetitive tabletop pick-and-place vs. bimanual manipulation with varied contacts.

Related event: Maxinsights Reports Data Scaling Laws for Physical AI(2 posts)→

Original post →

More from Embodied

Embodied channel →