New physics benchmark: top video model Seedance 2.5 scores just 57.76/100
机器之心 · wechat · 2026-10-08
NaversLab (EinsiaAI) released "World Models' Last Exam in Physics," testing 8 video generation models on 40 tasks across 9 physics phenomena (mechanics, optics, fluids, phase change, electromagnetism).
- Scoring: a consistency score C (VLM-judged coherence) plus a physics score P computed by programmatically measuring trajectories, periods, angles and liquid levels against predefined physical relations — no VLM subjective judgment. Physics counts 85% of the total and only counts if C≥80.
- Results: Seedance 2.5 tops the board at 57.76/100, MiniMax H3 follows at 54.89. On phase change, only Seedance passes (95.61) while other models max out at 15. Free-fall tasks average 98.75 consistency but just 13.88 physics — videos look smooth while motion violates physics. All models average under 40 on refraction/total internal reflection.
- Human alignment: with 10 annotators over 320 videos, the measurement-based method beats direct VLM scoring by 8.18/7.50 percentage points.
The takeaway: coherent-looking frames are far from physically correct futures, and temporal dynamics modeling remains the shared weakness of current video models.
Related event: New Physics Benchmark Tests Video World Models, Top Score Just 57.76(3 posts)→
More from Multimodal
- Higgsfield casts viral AI influencers in an AI-generated short drama — SimplyAnnisa · 2026-10-09
- First fully AI-generated film 'Gods Don't Give Gifts' gets official MPA rating, eyes Academy Awards run — Polymarket · 2026-10-09
- Google and Unity launch AI playground that builds games from text, plus Spark — glenbeer · 2026-10-09
- Hugging Face Speech-to-Speech adds Persian ASR and TTS support — andimarafioti · 2026-10-09
- ComfyUI app ANIMA goes viral: Qwen 3.5 'hallucinates' popular photos, 300K views — Ok_Contribution8157 · 2026-10-09
- NAMVIS (NeurIPS 2026): next-scale autoregression beats diffusion for novel-view synthesis, 3x faster — RexDouglass · 2026-10-09