Rhoda AI: scaling web-video pre-training keeps improving real-robot manipulation policies
stepjamUK · x · 2026-09-16
Rhoda AI rigorously tests the belief that scaling web-video pre-training improves downstream robot performance, using Direct Video-Action models (DVA) evaluated on a real industrial manipulation task across model sizes and compute budgets.
Key findings from 200+ hours of real-robot evaluation:
- Larger pre-trained video models consistently yield better robot policies, with no saturation at the largest size tested
- More pre-training compute helps at every amount of robot demonstration data, with the biggest gains when robot data is scarce
- Better held-out web-video prediction correlates with better real-robot policies across all scales
The results support scaling web-video pre-training as a core recipe for robot foundation models.
More from Embodied
- Robotics Researcher: The GPT of Robotics Will Just Be GPT — chris_j_paxton · 2026-09-16
- U-BOT waypoint-following MuJoCo sim looks 'too polished to be real' — _Stocko_ · 2026-09-16
- Humanoid Robot Shows Zero-Shot Navigation of Complex Terrain on Onboard Compute — chris_j_paxton · 2026-09-16
- Chinese garment factory workers wear head cameras to collect data for training humanoid robots — mustafamhus · 2026-09-16
- Meta reportedly plans camera-free Luna smart glasses this fall — Polymarket · 2026-09-16
- Ex-Tesla Optimus engineer says physical AGI still needs breakthroughs in video models and scaling laws — mihdalal · 2026-09-16