Q-Planning boosts robot self-improvement from 25% to 80% success
chris_j_paxton · x · 2026-08-26
Addressing the limitation of robot foundation models on hard tasks (e.g., 25% success rate), Q-Planning introduces a learning-based harness that enables self-improvement by equipping a large visuomotor BC policy with a small off-policy Q-function.
Key Mechanisms:
- Inference: The frozen BC policy samples N candidate action chunks; the Q-function scores them, and a single-step Q-weighted average is executed.
- Training: Successful and failed rollouts are appended to a replay buffer to fine-tune only the Q-function, leaving BC weights untouched.
Results:
- On LIBERO and bimanual RoboTwin benchmarks, ten iterations lifted all scores (e.g., LIBERO-10 93% → 99%, RoboTwin 83.8% → 91.4%).
- On two contact-rich real-robot tasks, the loop improved performance purely from autonomous interaction (no human intervention) in 30 minutes.
More from Embodied
- Chinese humanoid robot sales dwarf US counterparts — teortaxesTex · 2026-08-26
- Anchor-Align: Recovering OOD Generalization in VLA Fine-tuning — _krishna_murthy · 2026-08-26
- Chinese startup Rochu Robotics unveils hydraulic-driven biomimetic humanoid hand — teortaxesTex · 2026-08-26
- Why wheeled robots deploy easier than bipeds, per MIT expert — BradPorter_ · 2026-08-26
- First Robot Fight Betting Event: 200lb T800 Championship — zealcaiden · 2026-08-26
- SF robotics company builds custom mic array for better AI interaction — Scobleizer · 2026-08-26