Q-Planning Enables Self-Improving Robot Policies: LIBERO-10 Hits 99% With Frozen BC
chris_j_paxton · x · 2026-09-17
Q-Planning (Giridhar, Garg et al.) pairs a large visuomotor BC policy with a small off-policy Q-function that absorbs both successful and failed rollouts. Inference uses Q-weighted action selection; self-improvement fine-tunes only the Q-function, freezing BC weights. Ten iterations lift LIBERO-10 from 93% to 99% and bimanual RoboTwin from 83.8% to 91.4%, with similar gains on real contact-rich tasks—no human intervention. Code open-sourced.
More from Embodied
- Reka open-sources RekaDaily-10k: 10,000+ hours of real egocentric household robot training data — artetxem · 2026-09-17
- Grounded API launches with SOTA hand-tracking (<1cm) and SLAM benchmarks — databoydg · 2026-09-17
- Robotics team turns an actuator race condition bug into the feature they needed — eigenron · 2026-09-17
- ArmSoM Sige 7 unboxed: an RK3588 board for edge AI and robotics — chrismatthieu · 2026-09-17
- Closed-room Hong Kong robotics workshop tackles foundation models, data engines, deployment — paigeinsf · 2026-09-17
- Innate OS open-sourced: an agentic OS for general-purpose robots under $1k — ycombinator · 2026-09-17