WROP: 150 cognitive tasks train object permanence into video world models
Haotian Zhang · hf · 2026-09-25
Object permanence is a hallmark of human cognitive priors. Researchers introduce WROP (World Reasoning with Object Permanence), a core-cognition-inspired data infrastructure of 150 hand-designed cognitive science tasks across six categories. Blender generators randomize speed, lighting, and camera angle while preserving each task's cognitive structure, yielding 10,000+ samples per task.
They release a 1.5M-sample training corpus and a 300-question exam, then evaluate 14 video models (3 reference-to-video, 7 edit, 4 continuation) including their 16B world model PWM-WROP. In a blind pairwise Elo study, PWM-WROP ranks first among continuation models and third overall, behind a statistical tie of two reference-to-video models. Data, exam, scores, weights, and the native-PyTorch PWM training stack on AWS Trainium2 are all open-sourced.
More from Multimodal
- Midjourney's New Edit Model Keeps Characters Consistent Across Edits — tisch_eins · 2026-09-25
- Gemini Live translates English-to-Japanese sushi order in real time at Google's audio event — jacalulu · 2026-09-25
- Indie dev claims Opus 5.5 jumps AI video editing a full model generation — yihui_indie · 2026-09-25
- FLORA's multi-model creative workflow now works inside ChatGPT — round · 2026-09-25
- Wan 2.2 consistency experiment chains three 14s clips into one action scene — Tokyo_Jab · 2026-09-25
- H3 consistency experiment: character sheets plus tail-frame refs in one generation — Tokyo_Jab · 2026-09-25