Peking University's DexPolicy lifts dexterous manipulation success to 85% via annealed exploration
PekingUniversity · hf · 2026-10-02
DexPolicy from Peking University makes RL exploration scale an explicit function of training steps for trajectory-guided dexterous manipulation, annealing from broad to narrow exploration while keeping loss, architecture, reward, and optimizer fixed — addressing the conflict between noise needed for contact discovery and noise that hurts precise control.
- Compared PPO, critic-free GRPO continuation, and flow-parameterized PPO (FPO): on five YCB objects, mean deterministic Target success rose 49.4%→68.1% (FPO), 14.1%→45.4% (GRPO), 32.0%→35.7% (PPO);
- On a real RealMan RM75 arm with an Inspire/RH56 hand (360 trials), FPO lifted mean success from 25.0% to 85.0%, GRPO to 63.3%, PPO to 43.3%;
- Training return, deterministic success, and robustness to execution noise dissociate — schedules should be judged by terminal task success under intended execution conditions. Code and website are open-sourced.
More from Embodied
- YC F26 startup Preload captures synchronized video, EMG and tactile data for robots — ycombinator · 2026-10-02
- Robotics researcher: domain randomization trades away precision, classical control has better guarantees — KyleMorgenstein · 2026-10-02
- Qualcomm's IVD benchmark shows VLMs lag far behind humans at real-time real-world Q&A — rishit_dagli · 2026-10-02
- 'A robot without people to design and fix it is inventory, not revenue' — MatthewChang · 2026-10-02
- Galaxy General's ET1 humanoid starts at 79,000 yuan, learns dance moves by watching — 量子位 · 2026-10-02
- Robotics papers on arXiv hit 210 per day, up from 20-30 four years ago — DJiafei · 2026-10-02