Researchers teach a humanoid robot to learn from its own mistakes
imjustnewatai · x · 2026-08-04
Researchers from UC Berkeley, Google DeepMind, and NVIDIA showed that a humanoid robot can improve from its own mistakes without collecting new human demonstrations.
Using Gemini Robotics On-Device inside a Unitree G1, the team ran three real-world tasks and used a reward model to label which steps helped or hurt. Those interactions were converted into new training data for another round of fine-tuning.
Results after two practice cycles, across 100 trials per task:
- Box packing: 93% → 100%
- Cup insertion: 70% → 98%
- Plate handover: 53% → 96%
The key behavioral change was not just higher success rates: the robot started to rotate the box before grasping it and retried insertion after a failed first attempt, behaviors that were absent from the original human demonstrations.
![paper]
More from Embodied
- RethinkX says humanoid robots could deliver astronomical productivity returns — adam_dorr · 2026-08-04
- Opinion: AI Hardware Will Be Like Clothing, Multi-Device is the Future — msalbergo · 2026-08-04
- Surgical robots still give doctors less spatial awareness than a Tesla, founder says — ditzikow · 2026-08-04
- Using Codex, a beginner built a desk cat paw that taps them when they sit too long — 数字生命卡兹克 · 2026-08-04
- Raspberry Pi caption device turns calls and room talk into live subtitles — tom_doerr · 2026-08-04
- H3 video model reportedly fixes the gymnast-proportion problem and supports 3D motion reconstruction — andrew_n_carr · 2026-08-04