Robot AI today: VLA fast-slow stacks converge, on-bot continual learning is the missing piece

PTrubey · x · 2026-09-27

Convergent design

Figure's Helix, Physical Intelligence's π0 and Google DeepMind's Gemini Robotics all use vision-language-action models: a slow "thinking" layer (a multimodal LLM) understands requests, observes the scene and plans; a separate fast control layer issues motor commands many times per second. Bots can converse, decompose "clean up the kitchen" and retry failed grasps.

Training

Mostly imitation learning from human demos (teleoperation plus first/third-person video), with RL layered on top — locomotion mostly in simulation, fine skills increasingly fine-tuned in the real world.

Handling mistakes

What's missing

Context-based adaptation works for reasoning but not motor skill. A human fumbles twice with an unfamiliar tool and has it by the third try because their motor system rewires; current robots can't consolidate attempts into lasting skills until the next fleet retrain. Rare, local, site- or body-specific issues are handled badly (the author's Tesla keeps driving over the same road dip too fast).

Whether true on-bot continual learning arrives — and whether it matters — are the trillion-dollar questions in AI robotics today.

Original post →

More from Embodied

Embodied channel →