Q-Planning Enables Self-Improving Robot Policies

A team including Animesh Garg proposed Q-Planning, which pairs a frozen behavior-cloning policy with a small off-policy Q function for self-improvement, reaching 99% on LIBERO-10.

2026-09-17 ~ 2026-09-18 · 2 related posts