30,000 hours of ego-centric video: scaling data fixes agents, not object interactions

arankomatsuzaki · x · 2026-10-10

A CMU-led paper evaluates world models trained on 30,000 hours of ego-centric video (1,000+ scene types, 14,000 contributors), measuring agent and object-interaction fidelity directly on an out-of-distribution benchmark.

Bottom line: scaling ego-centric video nears the limit of agent modeling, while the ability to affect the world lags behind—closing the gap depends on how models are trained, not just how much data they see.

Original post →

More from Embodied

Embodied channel →