SOLE-R1 accepted at NeurIPS: zero-shot reward prediction teaches robots 20+ new tasks from scratch
___Mufasaa · x · 2026-09-30
- The SOLE-R1 team announced their paper on general-purpose reward models was accepted to NeurIPS 2026.
- Key result: SOLE-R1's zero-shot reward prediction serves as the sole reward signal for online RL, letting robots learn 20+ unseen manipulation tasks from randomly initialized policies (0% starting success rate).
- Implication: RL with zero-shot dense reward prediction can go beyond improving existing tasks — robots can acquire entirely new skills from zero competence, with no demonstrations or ground-truth rewards.
More from Embodied
- Berkeley's tactile compressor maps 10 fingertip streams to 2 hand latents, 2.26x faster training — berkeley_ai · 2026-09-30
- Berkeley's DexTacWAM turns video world models visuo-tactile, averaging 70.6 vs 38.0 on dexterous tasks — berkeley_ai · 2026-09-30
- Robot duck steals OpenAI DevDay: seamless control wows the crowd — gabrielchua · 2026-09-30
- IDC: China accounted for 77.9% of global humanoid robot shipments in H1 — yogthos · 2026-09-30
- Flourish home robot can be taught new tasks in 30 minutes via smartphone demos — philfung · 2026-09-30
- Friend founder pitches wearable AI pendant that could replace your phone and computer — jesselyu · 2026-09-30