How Google's RT-2 triggered the robotics boom: Understanding AI explains VLA models
binarybits · x · 2026-09-03
Timothy B. Lee's Understanding AI explainer (Robot Week) walks through vision-language-action models with minimal math. Key points: LLM researchers see GPT-3 (175B parameters, 300B tokens) as the real 2020 breakthrough, and it took years of long-context reasoning, tool use and context management to reach Claude Code in 2025. Robotics' "GPT-3 moment" came in July 2023 when Google's RT-2 trained a multimodal LLM to directly output robot actions — billions of parameters vs. 35M in predecessor RT-1 — kicking off today's robotics boom, which the author argues is on a similar multi-year path from base model to capable agents.
More from Embodied
- Counterfactual Debugging: causal attribution over 1M steps to localize sim2real gaps in world-model agents — MichaelD1729 · 2026-09-03
- Third-party test of MolmoAct 2: color-sensitive VLA that still struggles with small-object grasping — YuXiang_IRVL · 2026-09-03
- Agility Robotics' Digit graduates from fixed totes to handling arbitrary boxes — jonstephens85 · 2026-09-03
- Dev hooks busy bar up to Gumloop to show AI agents working live — aronkor · 2026-09-03
- Maker estimates home heating oil level with ESP32 and a clamp-on current sensor — JeremyCMorgan · 2026-09-03
- ARK: Data is the bottleneck for humanoids - Figure AI's app offers a fix — DMaguireARK · 2026-09-03