SOLE-R1: video-language reasoning as the sole reward for on-robot RL, at NeurIPS
micoolcho · x · 2026-10-09
MIT & RAI Institute's SOLE-R1 (NeurIPS 2026) uses a video-language reasoning model as the sole reward signal for online RL. Key points:
- Problem: VLMs as RL evaluators fail under partial observability and distribution shift, letting policies exploit perceptual errors.
- Method: per-timestep spatiotemporal CoT reasoning over raw video plus a natural-language goal yields dense progress estimates used directly as rewards; trained via a large-scale trajectory/reasoning synthesis pipeline and hybrid SFT + RL from verifiable rewards.
- Results: across 4 sim environments and real robots, zero-shot online RL from random init learns unseen manipulation tasks without ground-truth rewards, success indicators, demos, or per-task tuning.
Paper, code, models, and data are public; a RoboPapers episode is coming.
More from Embodied
- Mecka raises $60M Series B led by Sequoia to collect human motion data for training robots — emmanuelvivier · 2026-10-09
- London team preps five AI-powered Promptable Products, shipping Autumn 2026 — genmon · 2026-10-09
- Rohit Prasad starts first day as Boston Dynamics CEO, bets big on Physical AI — RoboBalaji · 2026-10-09
- VPP2: 14B Video Prediction Policy Beats Cosmos3-64B by 11 Points on Zero-Shot Manipulation — zhenjun_zhao · 2026-10-09
- Haiku 5.5 Hits 85% Success on Robot Tasks at Under $0.02 Per Attempt — scaling01 · 2026-10-09
- shadcn pitches a phone-powered two-handed robot: slide in your phone, it becomes your Grok bot — shadcn · 2026-10-09