Stanford's MobileVISTA fixes mobile manipulation's pose brittleness with generative data augmentation

jiajunwu_cs · x · 2026-10-08

Stanford University and Toyota Research Institute introduce MobileVISTA, a data generation framework for mobile manipulation. End-to-end policies trained on demos from a single base pose break when navigation leaves the base just a few centimeters off, pushing egocentric observations and end-effector trajectories out of distribution.

MobileVISTA jointly (1) augments egocentric visual observations with generative perturbations across poses and (2) retargets actions to compensate for base pose changes. Unlike prior work, it targets egocentric platforms like humanoids where the camera rides the actuated kinematic chain, and requires no extra demo collection or trained generative model.

Tested in simulation on humanoid and bimanual tasks and on a real Galaxea R1 Pro, policies trained on MobileVISTA-augmented data show markedly better robustness to previously out-of-distribution poses, with the largest gains on humanoids.

Original post →

More from Embodied

Embodied channel →