HG-DAgger Pitfall: Annealing With Clean Demonstrations Only Lifts Robot Task Success to 88%

DominiqueCAPaul · x · 2026-09-30

Robotics practitioner @aurelarnold shares a hands-on lesson from HG-DAgger (human-intervention imitation learning) training.

The problem: after a second round of intervention data collection, policy π2 barely improved on success rate and had a much longer mean episode duration. The culprit is intervention quality — when a policy gets into a bad state, human interventions are slower and less smooth than normal demonstrations because the operator is out of flow, so too many of these teach the policy to act slower and depend on recoveries.

The fix: finish the training run's annealing phase with clean demonstrations only. This mostly recovers fast, precise movement from the demonstrations while keeping the recovery skills learned from interventions.

Result: success rate on the "full box" task reached 88%. Directly useful for teams doing teleoperation/intervention-based robot policy training.

Original post →

More from Embodied

Embodied channel →