MIT's SOLE-R1: Video-Language Reasoning as the Sole Reward Enables Zero-Shot On-Robot RL
micoolcho · x · 2026-10-08
MIT and RAI Institute published SOLE-R1: Video-Language Reasoning as the Sole Reward for On-Robot RL (paper, code, models and data released).
Problem: When VLMs act as evaluators in RL, today's strongest models often fail under partial observability and distribution shift, letting policies exploit perceptual errors instead of solving the task.
Method: SOLE-R1 is a video-language reasoning model explicitly built to serve as the sole reward signal for online RL. Given only raw video and a natural-language goal, it performs per-timestep spatiotemporal chain-of-thought reasoning and outputs dense progress estimates used directly as rewards. Training uses a large-scale video-trajectory and reasoning synthesis pipeline producing temporally grounded CoT traces, combined with SFT and RL from verifiable rewards.
Results: Across 4 simulation environments and a real-robot setting, SOLE-R1 enables zero-shot online RL from random initialization — robots learn 24 unseen manipulation tasks with no ground-truth rewards, success indicators, demonstrations, or task-specific tuning.
More from Embodied
- Robotics Company Goes Viral Making a Music Video Starring Its Robots — adityaag · 2026-10-08
- Robotics is having its moment: why roboticists are Renaissance generalists — lukas_m_ziegler · 2026-10-07
- PhD student wraps up robotics planning series with EMPIRIC residual world model — tomssilver · 2026-10-07
- Wearing a BCI headband like Chinese quant traders is the new productivity hack — alexbilz · 2026-10-07
- Robot industry insider: no real-world 1X Neo deployments seen, Agility out of top 10 — jonstephens85 · 2026-10-07
- AI firm HUMXN offers free plumbing and HVAC service in Minnesota to train home robots — brunomarsjams · 2026-10-07