Amazon's StructRL uses verifiable intermediate rewards to improve long-horizon robot manipulation

amazon · hf · 2026-09-30

StructRL is an online RL framework for long-horizon vision-language-action tasks. Instead of sparse terminal rewards, it decomposes tasks into verifiable subtasks, grants intermediate rewards only after prerequisite subtasks complete, and scales rewards by completion pace. On RoboCasa365 and LIBERO-Long with GR00T-N1.5 and pi 0.5, StructRL consistently outperforms online RL baselines, showing structured intermediate rewards improve long-horizon VLA post-training. Code is open-sourced.

Original post →

More from Embodied

Embodied channel →