MIT's SOLE-R1: Video-Language Reasoning as the Sole Reward Enables Zero-Shot On-Robot RL

micoolcho · x · 2026-10-08

MIT and RAI Institute published SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot RL (paper, code, models and data released).

Problem: When VLMs act as evaluators in RL, today's strongest models often fail under partial observability and distribution shift, letting policies exploit perceptual errors instead of solving the task.

Method: SOLE-R1 is a video-language reasoning model explicitly built to serve as the sole reward signal for online RL. Given only raw video and a natural-language goal, it performs per-timestep spatiotemporal chain-of-thought reasoning and outputs dense progress estimates used directly as rewards. Training uses a large-scale video-trajectory and reasoning synthesis pipeline producing temporally grounded CoT traces, combined with SFT and RL from verifiable rewards.

Results: Across 4 simulation environments and a real-robot setting, SOLE-R1 enables zero-shot online RL from random initialization — robots learn 24 unseen manipulation tasks with no ground-truth rewards, success indicators, demonstrations, or task-specific tuning.

Original post →

More from Embodied

Embodied channel →