Principia benchmark: video models score 0.8 on VBench but under 0.42 on physics consistency
CSProfKGD · x · 2026-09-07
Researchers from IISc and Johns Hopkins introduce Principia, a benchmark testing whether video generation models understand high school physics via relational tests.
- Core idea: absolute motion measurements depend on frame rate, scale, and camera calibration, so Principia instead checks physical relationships between paired objects obeying the same law — calibration-independent by design.
- Coverage: eight phenomena (gravity, restitution, friction, rotational inertia, projectile motion, momentum, pendulum, mass-spring oscillation) across 500+ real controlled scenes.
- Results: five state-of-the-art video generators all score 0.8 on VBench yet none exceeds 0.42 on Principia — photorealistic motion doesn't mean physically correct motion.
- VLMs struggle too: the best vision-language model detects relational physics violations at only 67% accuracy; most perform near chance.
Paper and benchmark page are public, highlighting a large gap between visual realism and physical understanding in video generation.
More from Multimodal
- Turning AI-generated pixel sprites into animations with Sprite Fusion — HugoDuprez · 2026-09-07
- Grok, ChatGPT and Gemini All Failed at Making a Collage of His Book Covers — pickover · 2026-09-07
- Designer concedes GPT-6 Astra 'nearly unbeatable' after it designed full Figma pages via Computer Use — deedydas · 2026-09-07
- GPT-6 Astra Turns Van Gogh's Starry Night Into a Walkable 3D World — emmanuelvivier · 2026-09-07
- GPT-6 Astra's Iterative 3D Modeling Demos Impress Early Users — Cklly2004 · 2026-09-07
- AI Short Film 'Mechanical Love' Made with Seedance 2.5 and GPT Image 2.0 — HashemGhaili · 2026-09-07