VPP2: 14B Video Prediction Policy Beats Cosmos3-64B by 11 Points on Zero-Shot Manipulation
zhenjun_zhao · x · 2026-10-09
Video Prediction Policy 2 (VPP2): Predict Better, Act Better
- Problem: Existing world action models (WAMs) often produce wrong motion predictions in open-ended environments, because base video models aren't optimized for manipulation and naively adding action modules degrades generalization.
- Method:
- Curated a large, diverse manipulation-video dataset and performed event-level captioned video pretraining;
- Post-trained and distilled the video model into a single-step visual planner with fixed horizon;
- Added an action module via a mixture-of-transformers (MoT) architecture learning an implicit inverse dynamics model.
- Results: VPP2-14B outperforms Cosmos3-64B by 11.0 points, achieving strong zero-shot generalization in both video prediction and action generation.
More from Embodied
- Intel's John Healy: most enterprise robotics value sits in the installed base — ryanshrout · 2026-10-09
- IROS 2026 Best Paper: Feeding 3D Features Directly into VLM Lifts Navigation Success to 74.2% — qinzytech · 2026-10-09
- Saudi Robotics Firm Sirr Prices Its Exobot Humanoid at 399,000 Riyals — aziz4ai · 2026-10-09
- London team preps five AI-powered Promptable Products, shipping Autumn 2026 — genmon · 2026-10-09
- Rohit Prasad starts first day as Boston Dynamics CEO, bets big on Physical AI — RoboBalaji · 2026-10-09
- Haiku 5.5 Hits 85% Success on Robot Tasks at Under $0.02 Per Attempt — scaling01 · 2026-10-09