MolmoMotion: AI2's 4B VLM forecasts 3D point trajectories from language instructions
rsasaki0109 · x · 2026-10-03
MolmoMotion is a 4B vision-language model that forecasts 3D point trajectories under natural-language action instructions. Given a short RGB observation history, user-specified 2D query points with initial 3D positions, and a language description of the intended action, it predicts each point's 3D trajectory up to 2 seconds ahead in the camera-frame-at-t₀ coordinate system. The learned motion prior transfers to robotics planning and motion-guided video generation.
More from Embodied
- Self-sensing actuator senses forces down to 1g, 50x more precise than standard Chinese actuators — Scobleizer · 2026-10-03
- Open-source real-time SLAM plus hand reconstruction hits 30fps on a single Rockchip 3588 — micoolcho · 2026-10-03
- Meta Open-Sources Muse Code to Put AI in Your TV and Toaster — TechCrunch AI · 2026-10-03
- Calibrex: one-command sensor calibration checks for robot LiDAR, IMU, camera, GNSS — rsasaki0109 · 2026-10-03
- Video shows Waymo robotaxi navigating city streets and merging onto freeway — reed · 2026-10-03
- Watching My Home Robot Rooma Get Lost for the 1000th Time — OnlineInference · 2026-10-03