30,000 hours of ego-centric video: scaling data fixes agents, not object interactions
arankomatsuzaki · x · 2026-10-10
A CMU-led paper evaluates world models trained on 30,000 hours of ego-centric video (1,000+ scene types, 14,000 contributors), measuring agent and object-interaction fidelity directly on an out-of-distribution benchmark.
- 100x more data improves both, but unevenly: agent modeling is strong while object fidelity stays far lower and improves slowly.
- Careful visual conditioning saturates agent fidelity with a fraction of the data, enabling isolation of object fidelity and discovery of its saturation point.
- A supervision scheme shifting capacity toward object dynamics helps, but a substantial gap remains.
- Conclusions transfer to downstream humanoid modeling.
Bottom line: scaling ego-centric video nears the limit of agent modeling, while the ability to affect the world lags behind—closing the gap depends on how models are trained, not just how much data they see.
More from Embodied
- Starkey Unveils Omega AI+ Hearing Aids With G4 Gen AI Processor and Quad DNN — BrandonSawalich · 2026-10-10
- Wayve CEO demos AI Driver with Stellantis in a Fiat 500e on Turin's tight streets — alexgkendall · 2026-10-10
- 192GB Framework Desktop Batch 2 sells out; 128GB still in stock — film_girl · 2026-10-10
- Danu Robotics' six-year fight to build a better recycling robot — TechCrunch AI · 2026-10-10
- NVIDIA releases GR00T-based Agile One S SSD Pick robot deployment model on Hugging Face — _akhaliq · 2026-10-10
- Success-Guided Sampling: sim-to-real RL nails dexterous assembly with zero demos — kevin_zakka · 2026-10-10