HuRo: 630K robotized human-video episodes lift VLA completion from 51.5% to 80.3%
RLWRLD · hf · 2026-09-22
RLWRLD's HuRo pipeline converts heterogeneous human videos into robot-aligned observations and action trajectories, building a dataset of 630K robotized episodes and 142M frames from five human-video sources for VLA pretraining. Across four real-world manipulation tasks, more robotized pretraining data raises overall completion from 51.5% to 80.3% and OOD completion under spatial/visual shifts from 34.9% to 72.2%. Ablations show visual robotization improves OOD robustness and end-to-end retargeted-action pretraining beats visual-only transfer.
More from Embodied
- NUS's Grounded Action Model Tops Robot Manipulation Benchmarks with 3D Grounding — NationalUniversityofSingapore · 2026-09-22
- Distilling World-Model Features into VLAs: 0.8B Policy Hits 97.9% on LIBERO — Trung Dao · 2026-09-22
- HIRO Industries Unveils Origin: A 17-DOF Dual-Arm Robot for Packing and Kitting Workstations — Scobleizer · 2026-09-22
- Stealth team unveils new robot Origin; Scobleizer bets it would sell at Home Depot — Scobleizer · 2026-09-22
- Figure's Helix 2.5 robots tested on household tasks across 30 unseen homes — FinanceYF5 · 2026-09-22
- PrimeBOT T1 starts at $2,960 as line outputs a humanoid every 2.5 minutes — davidpattersonx · 2026-09-22