Alibaba's HappyWorld-Bench: 1,138 video cases test world model reliability
alibabagroup · hf · 2026-09-24
Alibaba released HappyWorld-Bench on Hugging Face, a comprehensive benchmark arguing world models must be judged not just on generation quality but on consistency and responsiveness as agents explore, interact with, and modify generated worlds.
- Framework: a hierarchical capability taxonomy (W1-W6) across three tracks — video, spatial, and embodied world models.
- Scale: 1,138 video prompts, 300 spatial scenes, 254 embodied cases; a HappyWorld-Arena runs human A/B comparisons producing model-level Elo ratings alongside automated behavioral-correctness metrics.
- Coverage: 14 video world models, 9 spatial systems, 8 embodied candidates evaluated.
- Findings: reliability gaps everywhere — video models lose consistency on long rollouts and revisits; spatial models top out at 70.14% placement accuracy and 73.33% edit execution; embodied models struggle to preserve state across multi-step actions and respond precisely to altered action conditions and physics.
More from Research
- Weaving directly shaped math's conceptual vocabulary, argues cognitive science thread — abenitezburraco · 2026-09-24
- Dual-H200 fine-tune of Marigold V2 with 16-frame temporal attention aims to fix video depth flicker — AntonObukhov1 · 2026-09-24
- Same Prompt, Opposite Results: GPT-4 Goes Silent 30/30 Where GPT-3.5 Never Stops — rayanpal_ · 2026-09-24
- Amazon's BoundaryMORPH uses Gaussian Processes to budget cross-encoder reranking, +5.4 nCG@100 — _reachsumit · 2026-09-24
- Paper: retrieval recall ceilings LLM recommendation reranking — oracle eval inflates NDCG up to 95%, none beat CF — _reachsumit · 2026-09-24
- Beyond a scalar: distributional serving interfaces let multiple task heads reuse watch-time distributions — _reachsumit · 2026-09-24