Q-Planning boosts robot self-improvement from 25% to 80% success

chris_j_paxton · x · 2026-08-26

Addressing the limitation of robot foundation models on hard tasks (e.g., 25% success rate), Q-Planning introduces a learning-based harness that enables self-improvement by equipping a large visuomotor BC policy with a small off-policy Q-function.

Key Mechanisms:

Results:

Original post →

More from Embodied

Embodied channel →