Robot walking demo adds target-facing rewards but still fails 80% of the time
carlosdponx · x · 2026-08-04
The author says they tried a better observation scheme for relative target location, added new rewards for facing the target and staying near it, and pre-conditioned the policy on walking motions with SFT.
The demo looks better at times, but it still reaches the target only about 20% of the time. They note that more compute may help and that the current setup is trained on 4090s, which only fit 128 parallel environments for this model.
More from Embodied
- ScienceCorp launches Scifi 2 headstage for brain recording and stimulation — gottapatchemall · 2026-08-04
- FCC ban on new foreign-made advanced robots explained in a new video — carlosdponx · 2026-08-04
- Instant re-inference and no interpolation lift a baseline policy from 76% to 88% — DominiqueCAPaul · 2026-08-04
- Seldon AI cofounder says robotics will hit jobs slower than model-driven automation — luke_drago_ · 2026-08-04
- A $30 humanoid robot cleaned an apartment better than a Roomba, says Seldon AI founder — luke_drago_ · 2026-08-04
- Aero Hand Open debuts as a $314 open-source robotic hand with 16 joints — TinfoilTricorn · 2026-08-04