Robotics data bottleneck: NVIDIA discards 96% of training video — search-first indexing as the fix
AI Engineer · youtube · 2026-09-24
Bright Data's Rafael Levi argues physical AI's next bottleneck is finding the right video: LLMs train on trillions of words, but robotics has only about a million videos of robots acting.
- Staged recordings produce biased data (nobody opens a door naturally on command); simulation and teleoperation don't scale
- The web holds billions of hours of real people handling objects with real physics
- Noise is costly: NVIDIA discards 96% of video downloaded for Cosmos, Stable Video Diffusion discards 74%
- Counterpoint: Meta trained on 1M hours of public video plus only 62 hours of real robot data to control a robot
Bright Data's "search first, collect second" approach indexes 1B+ videos by the actions in them, returning trimmed clips with timestamps, match scores and frame counts via API — applicable to self-driving and brand discovery too.
More from Embodied
- HRI 2027 Adds Archival Industry White Paper Track for Real-World Robot Deployments — petitegeek · 2026-09-25
- Google dev imagines orchestrating AI agents via smart glasses: tap their shoulder, watch their screen — jason_mayes · 2026-09-25
- Climbing analysis in 3D with iPhone LiDAR: open-source SAM 3.1 + ViTPose demos — dosco · 2026-09-25
- Menlo open-sources Asimov 1 humanoid locomotion policy and Isaac Lab training code — freelerobot · 2026-09-25
- Human-Robot Dialogue workshop at IROS 2026: NVIDIA, MIT, Georgia Tech speakers lined up — dhadfieldmenell · 2026-09-25
- AI glasses shipments up 263% in H1; Zhiyuan delivers 20,000th robot — 创业邦 · 2026-09-25