How Reflex trains humanoids: 3-stage RL pipeline from control to RGB-D grounding

DJiafei · x · 2026-10-09

Reflex trains in three stages: (1) reinforcement learning for whole-body catching, (2) inferring box dynamics from delayed, incomplete observations, and (3) learning to recover that dynamics representation from RGB-D history while keeping the controller fixed. This decomposition gives each capability a direct training signal and makes visual learning substantially more scalable.

Related event: Reflex Lets Unitree G1 Catch Tossed Boxes in About a Second(8 posts)→

Original post →

More from Embodied

Embodied channel →