Amazon's StructRL uses verifiable intermediate rewards to improve long-horizon robot manipulation
amazon · hf · 2026-09-30
StructRL is an online RL framework for long-horizon vision-language-action tasks. Instead of sparse terminal rewards, it decomposes tasks into verifiable subtasks, grants intermediate rewards only after prerequisite subtasks complete, and scales rewards by completion pace. On RoboCasa365 and LIBERO-Long with GR00T-N1.5 and pi 0.5, StructRL consistently outperforms online RL baselines, showing structured intermediate rewards improve long-horizon VLA post-training. Code is open-sourced.
More from Embodied
- AnyStep-WAM Cuts Denoising Steps by Up to 85% in World Action Models Without Losing Success Rate — Rui Wang · 2026-09-30
- Simify: training-free real-to-sim framework solves robot spatial reasoning in seconds — Ed__Johns · 2026-09-30
- ZeroBot trains robot manipulation from scratch in 119 seconds with 87% success using generative real2sim — Ed__Johns · 2026-09-30
- LadderMan, zero-shot sim-to-real humanoid ladder climbing, wins CoRL 2026 Spotlight — yuewang314 · 2026-09-30
- Reddit users warm to Goose, a non-humanoid home robot that fetches and grabs — fhpapa · 2026-09-30
- Memory price surge threatens on-device AI hardware: BOM tops $300, device may cost $1,500 — stephen280up · 2026-09-30