CMU's R³ Trains VLMs to Reason in Natural Language to Guide Robots via RL
aviral_kumar2 · x · 2026-09-03
CMU (Aviral Kumar's group) released R³, studying whether VLMs can reason in natural language to guide low-level manipulation policies.
The recipe: mid-train a VLM on expert-generated reasoning traces to initialize the reasoning style, then improve with single-step rubric-based RL from offline action data. On Language Table and simulated bimanual grocery packing, R³ generalizes better on unseen tasks and significantly beats instruction-only imitation learning.
Key lessons from the authors: the basics (mid-training + rubric RL) done right can work; carefully choosing which tasks the VLM trains its own reasoning on matters.
Related event: CMU's R³ Teaches Robots to Reason Before Acting via RL(3 posts)→
More from Embodied
- Second Wayve robotaxi ride in the US suggests its driver is highly scalable — Sethwinterroth · 2026-09-03
- Trained a microduck to limbo: a charming tiny-robot demo — kevin_zakka · 2026-09-03
- Unitree founder Wang Xingxing says robotics' ChatGPT moment is still 2-3 years away — pstAsiatech · 2026-09-03
- Inside Dexory: what it takes to run a vertically integrated robotics company in Europe — lukas_m_ziegler · 2026-09-03
- Skild AI's S1 robot learns plant repotting, coffee and cooking from one video demo — Matt Wolfe · 2026-09-03
- Robot duck learns to skateboard on a Blender-MCP-designed 3D-printable board — TinfoilTricorn · 2026-09-03