Q-Planning adds a Q-function to policies like pi-0.5 so robots improve without interventions

micoolcho · x · 2026-09-18

On RoboPapers, Varun Giridhar3 and Animesh Garg present Q-Planning: imitation learning with interventions drives recent robotics progress, but fixing a policy via targeted interventions is labor-intensive. Q-Planning starts from a large policy like pi-0.5, adds a Q-function estimator for value prediction, updates it online using both successful and failed rollouts, and uses it to guide sampling and trajectory selection — dramatically improving performance with just a few rollouts, no human interventions needed.

Related event: Q-Planning Enables Self-Improving Robot Policies via Small Q-Functions(3 posts)→

Original post →

More from Embodied

Embodied channel →