MolmoMotion: AI2's 4B VLM forecasts 3D point trajectories from language instructions

rsasaki0109 · x · 2026-10-03

MolmoMotion is a 4B vision-language model that forecasts 3D point trajectories under natural-language action instructions. Given a short RGB observation history, user-specified 2D query points with initial 3D positions, and a language description of the intended action, it predicts each point's 3D trajectory up to 2 seconds ahead in the camera-frame-at-t₀ coordinate system. The learned motion prior transfers to robotics planning and motion-guided video generation.

Related event: AI2 Open-Sources MolmoMotion for Language-Guided 3D Point Trajectory Prediction(3 posts)→

Original post →

More from Embodied

Embodied channel →