Q-Planning: robot policies self-improve from 25% to 80% success, no new demos

rbhar90 · x · 2026-08-27

The demo-to-deployment gap in robotics is a data problem: the thread argues additional teleop SFT data on robot foundation models suffers from covariate shift, and ICL alone doesn't reach mastery. Q-Planning offers an alternative via test-time self-improvement.

Essentially a "thinking mode" for robot policies via test-time compute.

Related event: Q-Planning Enables Robot Self-Improvement, Boosting Success from 25% to 80%(2 posts)→

Original post →

More from Embodied

Embodied channel →