Researchers teach a humanoid robot to learn from its own mistakes

imjustnewatai · x · 2026-08-04

Researchers from UC Berkeley, Google DeepMind, and NVIDIA showed that a humanoid robot can improve from its own mistakes without collecting new human demonstrations.

Using Gemini Robotics On-Device inside a Unitree G1, the team ran three real-world tasks and used a reward model to label which steps helped or hurt. Those interactions were converted into new training data for another round of fine-tuning.

Results after two practice cycles, across 100 trials per task:

The key behavioral change was not just higher success rates: the robot started to rotate the box before grasping it and retried insertion after a failed first attempt, behaviors that were absent from the original human demonstrations.

![paper]

Original post →

More from Embodied

Embodied channel →