SGS tweak to RL resets lets sim-trained robots mesh gears at 94% zero-shot
abhishekunique7 · x · 2026-10-10
A University of Washington + NVIDIA team (CoRL 2026) introduce Success-Guided Sampling (SGS): the bottleneck in dexterous manipulation isn't a better RL algorithm but which task configurations environments reset to during training.
- SGS resets environments to configurations neither too easy nor too hard for the current policy — a small outer loop on standard PPO, with no demonstrations and no per-task reward tuning.
- Trained fully in simulation, policies transfer zero-shot to a real UR5e from RGB only: 94% success at gear meshing, plus M16 nut threading and 16mm peg insertion on the NIST task board.
- The same recipe handles legged locomotion: one policy takes ANYmal C/D across stepping stones, floating islands, and more.
- SGS keeps scaling past a million parallel simulated robots where uniform sampling plateaus. Paper is out; code coming soon.
More from Embodied
- Former SpaceX engineers raise $100M for self-driving electric freight trains — chrisgrayson · 2026-10-10
- Kleiner Perkins' Ilya Fushman: the moment for robotics is finally now — Sethwinterroth · 2026-10-10
- Dex-One2Many: Real2Sim2Real Turns One Human Video Into Robots Generalizing Across Configurations — furongh · 2026-10-10
- Coevolved robot communication transfers poorly to 3D: 1 success in 30 seeds — uv-mex · 2026-10-10
- Transferring co-evolved robot communication from 2D to 3D physics simulation — uv-mex · 2026-10-10
- Blogger suggests adding CUA data to robotic model pre-training mix — eigenron · 2026-10-10