MBZUAI's WorldGuide beats MiniMax-H3 on closed-loop procedural video world modeling
MBZUAI · hf · 2026-10-09
- MBZUAI introduces WorldGuide, formulating procedural video generation as closed-loop task execution in visual world space: given only an initial image and goal, it predicts atomic actions, generates corresponding video clips, and uses the results to pick next actions or terminate.
- Planner and Executor are trained on the same step-level demonstrations—the Planner predicts the next atomic action or completion from visual progress, the Executor learns to realize predicted actions—with hierarchical visual memory maintaining long-horizon state at bounded token cost.
- Lacking step-level action-video supervision, the team built WorldGuide Bench (59K step-annotated videos, 245 tasks, 27 categories). WorldGuide hits 33.33% task success vs 29.90% for MiniMax-H3 (which even gets reference action plans), and 47.69% vs 32.73% on VideoCraft-Bench under goal-only conditioning.
More from Research
- How Those Morphogenesis Simulations Work: Deformable Meshes With Controllable Fibers — zzznah · 2026-10-09
- OpenAI's frontier model produces new math results, including an 84-page proof of the Erdős–Pomerance conjecture — burny_tech · 2026-10-09
- Researcher: activation monitors beat black-box monitoring for AI cyber safety — burny_tech · 2026-10-09
- ETH's SpaceFlow: training-free locally controllable 3D generation from text and primitives — ethz · 2026-10-09
- EDiS: cached edge-disjoint subgraphs cut GNN sparse-training cost and top benchmarks — Sai Karthik Navuluru · 2026-10-09
- Yi Ma and Yann LeCun spar over automation's coming reshaping of mathematics — CSProfKGD · 2026-10-09