Q-Planning adds a Q-function to policies like pi-0.5 so robots improve without interventions
micoolcho · x · 2026-09-18
On RoboPapers, Varun Giridhar3 and Animesh Garg present Q-Planning: imitation learning with interventions drives recent robotics progress, but fixing a policy via targeted interventions is labor-intensive. Q-Planning starts from a large policy like pi-0.5, adds a Q-function estimator for value prediction, updates it online using both successful and failed rollouts, and uses it to guide sampling and trajectory selection — dramatically improving performance with just a few rollouts, no human interventions needed.
Related event: Q-Planning Enables Self-Improving Robot Policies via Small Q-Functions(3 posts)→
More from Embodied
- Is McDonald's the best-positioned restaurant chain for robots — and is it too early? — clemnt · 2026-09-18
- STM32 robotic hand packs tendon-driven fingers, FSR and Hall sensors into five digits — _Stocko_ · 2026-09-18
- Tesla FSD Supervised now live in 14 markets across four continents — XFreeze · 2026-09-18
- Muse model is out driving Frodobots' Earth Rover Mini robot — micoolcho · 2026-09-18
- Chris Paxton: Nobody Wants a $20,000 Home Robot That Breaks 1% of Their Dishes — chris_j_paxton · 2026-09-18
- KUKA CEO Talks Physical AI at Augsburg HQ After 600,000 Robots Deployed — lukas_m_ziegler · 2026-09-18