R³ paper: robots learn to think before acting with carefully done RL on explaining demo data
aviral_kumar2 · x · 2026-09-03
The new R³ paper teaches robots to think before acting — free-form, deliberate reasoning that actually steers low-level policies.
How it works:
- Robot VLMs improve their own reasoning by practicing to explain demonstration data
- The recipe is surprisingly simple: RL, done carefully (mid-training + rubric-based RL)
- Trained reasoning VLMs then control low-level VLA policies
Why it matters:
- Language reasoning offers a cheap axis for scaling test-time compute for robots, enabling fast autonomous improvement and generalization
- Many SOTA robot learning methods still handcraft reasoning annotations
Lessons: task selection matters (the team ran controlled studies on when reasoning helps), and stronger base reasoning VLMs are critical.
Related event: CMU's R³ Teaches Robots to Reason Before Acting via RL(3 posts)→
More from Embodied
- Qwen-RobotManip: alignment before scale for robotic manipulation foundation models — rsasaki0109 · 2026-09-03
- WRC 2026 takeaways: humanoids pivot to industry solutions, tactile dexterous hands everywhere — CyberRobooo · 2026-09-03
- AM-ARM200: open-source 3D-printable 6+1 DoF robot arm with 1kg payload for ~$380 — RemiCadene · 2026-09-03
- EU rules 2022/1426 already cover fully driverless approval, but Ireland still lacks deployment rules for Cybercab — Graham_dePenros · 2026-09-03
- Nvidia brings RTX Spark AI PCs to Europe, blurring the line between chip and studio — nordicinst · 2026-09-03
- Figure commits $3.5B to deploy up to 100,000 NVIDIA Vera Rubin GPUs with Nscale — adcock_brett · 2026-09-03