Q-Planning: frozen BC policy plus small Q-function enables robot self-improvement

animesh_garg · x · 2026-09-18

The Q-Planning paper (Garg lab, podcast coming soon) tackles a core BC limitation: policies can't self-improve from failures without new human demos, while RL fine-tuning doesn't scale to billion-parameter visuomotor policies.

Related event: Q-Planning Enables Self-Improving Robot Policies(2 posts)→

Original post →

More from Embodied

Embodied channel →