Robot walking demo adds target-facing rewards but still fails 80% of the time

carlosdponx · x · 2026-08-04

The author says they tried a better observation scheme for relative target location, added new rewards for facing the target and staying near it, and pre-conditioned the policy on walking motions with SFT.

The demo looks better at times, but it still reaches the target only about 20% of the time. They note that more compute may help and that the current setup is trained on 4090s, which only fit 128 parallel environments for this model.

Original post →

More from Embodied

Embodied channel →