Sergey Levine on robotics: hardware is good enough, the real gap is decision-making and data
机器之心 · wechat · 2026-10-05
- Figure's Index data platform hit 16M+ human task videos from 108 countries in four months; Figure paid $15M and plans $1B+ in data/compute over the next year.
- Physical Intelligence co-founder Sergey Levine argues:
- Hardware is now good enough (Boston Dynamics was about control, not decisions); the frontier is the decision loop — AI must understand the world beyond the robot's body.
- Priors matter: Google's "arm farm" showed tabula-rasa learning plateaus at grasping; robots need priors from observation and experience, like humans assembling IKEA furniture who've seen furniture before.
- Counterintuitive data ordering: ground models in real embodied experience first, then they absorb simulation and internet video better — opposite of the popular YouTube-first route.
- Reliability is existential: unlike LLMs, robots must complete tasks autonomously to be useful; RL is key to going from 95% to 100% success, still unsolved.
- Ecosystem and China: China shipped 40k+ humanoids (97% of global) in H1, but the whole field has only 100k hours of embodied data; no single lab can build a foundation model alone — data must be shared.
- Robotics hasn't entered its stable scaling era yet; the industry must first find what's genuinely worth scaling, with Wang Xingxing estimating a "ChatGPT moment" is 2-5+ years away.
More from Embodied
- Dream4ACT Unifies Video-Action Modeling Across Robot Embodiments, Hitting 89% on RoboTwin 2.0 — Xiangyu Zhu · 2026-10-05
- Munich Robotics Startup RobCo Hits $1B Valuation in Nine Months — lukas_m_ziegler · 2026-10-05
- Physical AI field session in Bengaluru to tackle post-deployment evals and retraining loops — carrycooldude · 2026-10-05
- Meta's NAVA-WAM pretrains robot action policies directly from action-free videos — meta · 2026-10-05
- Robotics' most-hyped model completed just 7 of 100 real manipulation tasks, demo reel hides the rest — carrycooldude · 2026-10-05
- SJTU's LIFT adds force sensing to VLAs with zero force-labeled pretraining data — jiqizhixin · 2026-10-05