MotionVLA accepted at CoRL 2026: remember motion, not frames
yshan2u · x · 2026-09-08
MotionVLA, from Xinggang Wang's team, is accepted at CoRL 2026 (arXiv:2606.08288).
Core claim: VLA models should remember the motion connecting past frames, not the frames themselves. Existing approaches condition on history, depth, or 4D features to resolve long-horizon ambiguity, but motion-inconsistent evidence introduces geometric drift, fragmented temporal cues, and unstable action generation.
Method:
- Converts a short past-only video window into compact, time-continuous trajectory-field tokens;
- Current visual tokens query this history to retrieve task-relevant motion information;
- The retrieved motion is recoupled into the VLA stream under trajectory-grounded supervision.
Across simulation benchmarks and preliminary real-robot rollouts, MotionVLA improves long-horizon manipulation with smoother, more direct executions — effective VLA memory is about usable motion-consistent evidence, not more 4D context.
More from Embodied
- Dev Builds a TikTok-Style Mobile UI for 15+ Years of Replication Archives — Josikinz · 2026-09-08
- USC lab's SIMPLE lands at CoRL 2026: 60-task humanoid sim benchmark for VLA models — zhengyiluo · 2026-09-08
- Swarm Robotics Paper Tests When Sociality Pays Off: Scarce Environments Erase the Advantage — abenitezburraco · 2026-09-08
- GPT-6 Astra builds 3D site exploding a humanoid robot into 1,168 CAD parts — freelerobot · 2026-09-08
- Robotics researcher: harness + tool calls beat raw VLA on control benchmarks, at 2.3x lower cost — rbhar90 · 2026-09-08
- Grok bot plays a Roland piano, composing romantic-style music from a text prompt — yunta_tsai · 2026-09-08