Q-Planning Enables Self-Improving Robot Policies
A team including Animesh Garg proposed Q-Planning, which pairs a frozen behavior-cloning policy with a small off-policy Q function for self-improvement, reaching 99% on LIBERO-10.
2026-09-17 ~ 2026-09-18 · 2 related posts
- Q-Planning Enables Self-Improving Robot Policies: LIBERO-10 Hits 99% With Frozen BC — chris_j_paxton · 2026-09-17
- Q-Planning: frozen BC policy plus small Q-function enables robot self-improvement — animesh_garg · 2026-09-18