Dex-One2Many: Real2Sim2Real Turns One Human Video Into Robots Generalizing Across Configurations

furongh · x · 2026-10-10

Dex-One2Many uses a Real2Sim2Real pipeline with neuro-symbolic representations and a new technique called 'controlled diversification': a vision-language model converts one human demonstration video into scene-graph sequences capturing task structure, while the symbolic constraints specify what must hold and a neural policy learns how to achieve it. By generating diverse simulated training experiences that preserve task constraints, the resulting robot generalizes beyond the demo to new object positions, goal poses, and grasps.

Related event: Dex-One2Many: One Human Video Trains Multi-Morphology Robot Policies(2 posts)→

Original post →

More from Embodied

Embodied channel →