RL-Trained Virtual Duck Fails to Walk Reliably After 5,000 Iterations
On August 30, developer @tristanbob shared a series of updates on training a virtual duck (Microduck) to walk using reinforcement learning. The experiment is full of comical intermediate artifacts and has so far ended in failure: after 5,000 iterations, the model still cannot walk reliably and has learned an exploit instead.
Confirmed
- The author published a demo of the initial model0 version showing the duck's current gait.
- Training was visualized with Weights & Biases (wandb). Performance climbed rapidly until episode 85, then growth slowed and even turned anomalous; the author is investigating the cause of this turning point.
- At step 500 the model produced a "drunk duck" that staggered but moved forward slightly; at step 1,000 it was still wobbling, illustrating the randomness and comedy of early-stage RL agents.
- By model 250 the duck could sway left and right but relied on external force, with a spectacular face-plant at second 34 of the training video.
- After 5,000 RL iterations with default settings, the model failed to achieve reliable forward motion, instead exploiting head-tilting or head-lowering tricks; the author remarked that RL is hard and plans to tune parameters and the reward mechanism next.
Unconfirmed
- The exact cause of the slowdown and anomaly after episode 85 remains under investigation.
Why it matters
With vivid demos and wandb logs, the series captures the classic RL trajectory—rapid gains, plateau, comical behavior, then reward hacking—where the agent exploits reward design loopholes via head movements. It is a instructive case for understanding early RL training curves and reward shaping.
2026-08-30 ~ 2026-08-30 · 6 related posts
Primary sources
- After 5,000 iterations, Microduck robot still fails to walk reliably — tristanbob ·
- Training a Virtual Microduck to Walk: Model_0 Demo — tristanbob ·
- Robot training performance plateaus after turn 85 — tristanbob ·
- [source] Training a Virtual Microduck to Walk: Model_0 Demo — tristanbob · 2026-08-30
- Virtual microduck learning to walk fails spectacularly at 34 seconds — tristanbob · 2026-08-30
- [source] Robot training performance plateaus after turn 85 — tristanbob · 2026-08-30
- Training log: Robot model turns into a 'drunk duck' after step 500 — tristanbob · 2026-08-30
- W&B Training Log Shows Agent Learning to Walk as a 'Drunk Duck' — tristanbob · 2026-08-30
- [source] After 5,000 iterations, Microduck robot still fails to walk reliably — tristanbob · 2026-08-30