CMU's R³ Trains VLMs to Reason in Natural Language to Guide Robots via RL

aviral_kumar2 · x · 2026-09-03

CMU (Aviral Kumar's group) released R³, studying whether VLMs can reason in natural language to guide low-level manipulation policies.

The recipe: mid-train a VLM on expert-generated reasoning traces to initialize the reasoning style, then improve with single-step rubric-based RL from offline action data. On Language Table and simulated bimanual grocery packing, R³ generalizes better on unseen tasks and significantly beats instruction-only imitation learning.

Key lessons from the authors: the basics (mid-training + rubric RL) done right can work; carefully choosing which tasks the VLM trains its own reasoning on matters.

Related event: CMU's R³ Teaches Robots to Reason Before Acting via RL(3 posts)→

Original post →

More from Embodied

Embodied channel →