Peking University's DexPolicy lifts dexterous manipulation success to 85% via annealed exploration

PekingUniversity · hf · 2026-10-02

DexPolicy from Peking University makes RL exploration scale an explicit function of training steps for trajectory-guided dexterous manipulation, annealing from broad to narrow exploration while keeping loss, architecture, reward, and optimizer fixed — addressing the conflict between noise needed for contact discovery and noise that hurts precise control.

Original post →

More from Embodied

Embodied channel →