Q-Planning: frozen BC policy plus small Q-function enables robot self-improvement
animesh_garg · x · 2026-09-18
The Q-Planning paper (Garg lab, podcast coming soon) tackles a core BC limitation: policies can't self-improve from failures without new human demos, while RL fine-tuning doesn't scale to billion-parameter visuomotor policies.
- Attach a small off-policy Q-function to a large frozen BC policy; Q estimates value, so it absorbs both successful and failed deployment rollouts
- Inference: Q-weighted single-step average over BC action-chunk draws; self-improvement fine-tunes only Q, BC weights untouched
- Results: 10 iterations lift LIBERO-10 from 93% to 99% and RoboTwin from 83.8% to 91.4%; the same loop improves on two contact-rich bimanual real-robot tasks with zero human intervention
Related event: Q-Planning Enables Self-Improving Robot Policies(2 posts)→
More from Embodied
- GPT-6 Astra beats Claude Fable 5.1 on robot control: 35% vs 15% success — k7agar · 2026-09-18
- Doosan Robotics partners with peaq on self-monetizing robots, first unit already on peaqOS — LexSokolin · 2026-09-18
- Neuralink patient who lost speech says "I love you" via brain implant in his synthesized voice — Polymarket · 2026-09-18
- OpenHarness launches: open-source software + hardware to run Claude Code and Codex as CAD, PCB and robotics specialists — dee_hw · 2026-09-18
- NewEyes opens third-party access to Meta AI glasses, bringing sight-aware AI chat — rohanpaul_ai · 2026-09-18
- Collov Labs Launches NewEyes, a Visual AI Agent for Meta Glasses With Hybrid Edge-Cloud Design — FellMentKE · 2026-09-18