Q-Planning method lets robots self-improve, boosting success rates from 25% to 80%
animesh_garg · x · 2026-08-31
Traditional imitation learning struggles with precise tasks. The new Q-Planning method adds a small "critic" module to score potential next moves.
Core Mechanism:
- Keeps the original Behavior Cloning (BC) policy frozen.
- Updates only the Q-function, learning from both success and failure.
- Uses Q-weighted averaging for action selection during inference.
Results:
- In real-world bimanual robot tests, wallet insertion success rate rose from 25% to 80%.
- Cup-stacking tasks also saw significant gains.
- Performance improves in 30 minutes of autonomous attempts without extra human demos.
Paper: Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning
More from Embodied
- China dominates with 86% of global humanoid robot shipments in H1 2026 — rohanpaul_ai · 2026-08-31
- Experts Discuss Biomechanical Differences in Robotics — dhadfieldmenell · 2026-08-31
- How to Prepare for Tesla Cybercab Launch This Week — JOBhakdi · 2026-08-31
- Chinese Companies Unleash AI-Powered Robo-Chefs — yogthos · 2026-08-31
- S1's ICL advantage grows exponentially in long-horizon, out-of-distribution tasks — ZeYanjie · 2026-08-31
- Robot control trained in under 2 minutes on one 4090 via Sim2Sim transfer — yacineMTB · 2026-08-31