EgoHumanoid: Training Humanoid Robots via Egocentric Human Videos
micoolcho · x · 2026-08-06
Researchers have introduced EgoHumanoid, the first framework to co-train a vision-language-action (VLA) policy using abundant egocentric human demonstrations alongside limited robot data for whole-body loco-manipulation.
Methodology: To bridge the embodiment gap between humans and robots, the team designed a systematic alignment pipeline:
- View Alignment: Reduces visual discrepancies.
- Action Alignment: Maps human motions into a unified action space for humanoid control.
Results: Real-world experiments demonstrate that incorporating robot-free egocentric data significantly outperforms robot-only baselines by 51% in generalization performance, particularly in unseen environments.
More from Embodied
- Karpathy's AutoResearch Aims to Break Boundaries in Physical AI — teortaxesTex · 2026-08-06
- UMD Introduces HumanEgo: Training Robots via Human Demonstration Videos — furongh · 2026-08-06
- Fei-Fei Li on Spatial Intelligence: WorldLabs Acquires SceniX to Build Robotic Digital Training Grounds — 量子位 · 2026-08-06
- Open-Source DIY Translator Built on Raspberry Pi and Gemma — osanseviero · 2026-08-06
- Minimax video model runs locally: 4080 generates 1280x720 in ~30 mins — cocktailpeanut · 2026-08-06
- Developer working on reading experience for book scanner — aaraalto · 2026-08-06