MotionVLA accepted at CoRL 2026: remember motion, not frames

yshan2u · x · 2026-09-08

MotionVLA, from Xinggang Wang's team, is accepted at CoRL 2026 (arXiv:2606.08288).

Core claim: VLA models should remember the motion connecting past frames, not the frames themselves. Existing approaches condition on history, depth, or 4D features to resolve long-horizon ambiguity, but motion-inconsistent evidence introduces geometric drift, fragmented temporal cues, and unstable action generation.

Method:

Across simulation benchmarks and preliminary real-robot rollouts, MotionVLA improves long-horizon manipulation with smoother, more direct executions — effective VLA memory is about usable motion-consistent evidence, not more 4D context.

Original post →

More from Embodied

Embodied channel →