HG-DAgger Pitfall: Annealing With Clean Demonstrations Only Lifts Robot Task Success to 88%
DominiqueCAPaul · x · 2026-09-30
Robotics practitioner @aurelarnold shares a hands-on lesson from HG-DAgger (human-intervention imitation learning) training.
The problem: after a second round of intervention data collection, policy π2 barely improved on success rate and had a much longer mean episode duration. The culprit is intervention quality — when a policy gets into a bad state, human interventions are slower and less smooth than normal demonstrations because the operator is out of flow, so too many of these teach the policy to act slower and depend on recoveries.
The fix: finish the training run's annealing phase with clean demonstrations only. This mostly recovers fast, precise movement from the demonstrations while keeping the recovery skills learned from interventions.
Result: success rate on the "full box" task reached 88%. Directly useful for teams doing teleoperation/intervention-based robot policy training.
More from Embodied
- Hacky DIY leader arm claims 1/10th the cost of the original leader arm — k7agar · 2026-09-30
- Investor explains backing Figure AI at ~$2B: betting on $1T+ outcomes — markjeffrey · 2026-09-30
- Dyna Robotics ships Dyna-2.1, pivoting from task mastery to 'whole employee' robots paid by role — JasonMa2020 · 2026-09-30
- Wayve supervised ride in San Francisco and highway completed with zero interventions — alexgkendall · 2026-09-30
- Raspberry Pi's $30 Smart Display Module Ships With Optional Hailo AI Acceleration — JeremyCMorgan · 2026-09-30
- Dyna Robotics' Taku masters new whole-body tasks from just 30 minutes of data — charlieholtz · 2026-09-30