SOLE-R1: video-language reasoning as the sole reward for on-robot RL, at NeurIPS

micoolcho · x · 2026-10-09

MIT & RAI Institute's SOLE-R1 (NeurIPS 2026) uses a video-language reasoning model as the sole reward signal for online RL. Key points:

Paper, code, models, and data are public; a RoboPapers episode is coming.

Related event: MIT's SOLE-R1 Uses Video-Language Reasoning as Sole Reward for Robot Learning(2 posts)→

Original post →

More from Embodied

Embodied channel →