WorldExam benchmarks whether video models generate worlds that actually react
Yuxue Yang · hf · 2026-08-04
- WorldExam is a benchmark for treating controllable video generation systems as world models, going beyond visual quality to test whether the generated world reacts plausibly.
- It defines four diagnostic levels: Visual Quality, Control Adherence, Spatial Consistency, and World Reactivity.
- The benchmark spans 1,474 cases across 8 tasks and supports camera-, action-, and language-driven paradigms.
- Evaluating 20 representative models reveals a split: camera-driven models handle camera control well but lack dynamic interaction, action-driven models control subjects more precisely but often leave the world unresponsive, and language-driven models interact better but follow complex controls less faithfully.
- The takeaway is that good-looking videos and explicit instruction following are not enough to make a strong world model.
More from Multimodal
- Minimax H3 Omni uses After Effects motion references for image animation — bdsqlsz · 2026-08-04
- Multiple LLMs still fail to identify a Cubana Il-96 in a simple plane photo — airbus_a360_when · 2026-08-04
- ChatGPT now draws a watch showing the exact time you ask for — binary-baba · 2026-08-04
- A 10-second Pixar-style animation took 13 minutes on an RTX 3060 12GB — Pitiful_Archer_4381 · 2026-08-04
- MiniMax H3’s latest demo is funny, flawed, and still impressive — comfyui_user_999 · 2026-08-04
- A user is looking for MiniMax workflows that improve generation speed — PersonalMango2562 · 2026-08-04