Q-Planning Enables Self-Improving Robot Policies: LIBERO-10 Hits 99% With Frozen BC

chris_j_paxton · x · 2026-09-17

Q-Planning (Giridhar, Garg et al.) pairs a large visuomotor BC policy with a small off-policy Q-function that absorbs both successful and failed rollouts. Inference uses Q-weighted action selection; self-improvement fine-tunes only the Q-function, freezing BC weights. Ten iterations lift LIBERO-10 from 93% to 99% and bimanual RoboTwin from 83.8% to 91.4%, with similar gains on real contact-rich tasks—no human intervention. Code open-sourced.

Original post →

More from Embodied

Embodied channel →