Principia benchmark: top video models score ~0.8 on VBench but under 0.42 on physical consistency
anand_bhattad · x · 2026-09-05
The Principia paper introduces a benchmark testing whether video models obey Newtonian physics via relational invariants between paired objects, sidestepping frame-rate, scale, and camera-calibration ambiguity. It spans eight phenomena—gravity, restitution, friction, rotational inertia, projectile motion, momentum, pendulum, and mass-spring oscillation—using real-world scenes under controlled protocols.
Key findings:
- Across thousands of generations from six state-of-the-art video generators, all score 0.8 on VBench yet none exceeds 0.42 on physical consistency.
- The authors propose a calibration-independent consistency score that quantifies physical violations directly in image space.
- VLMs asked to detect relational physics violations top out at 67% accuracy, with most near chance level.
Related event: Principia Benchmark Exposes Poor Physics Reasoning in Video Models(3 posts)→
More from Multimodal
- Clapper hooks up fal's MiniMax H3 Director as flngr teases something new — flngr · 2026-09-06
- Runway becomes a featured Creativity plugin on Vercel — tlakomy · 2026-09-06
- Seedance 2.5 generates ultra-realistic 30s early-2000s MiniDV home video — SimplyAnnisa · 2026-09-06
- MiniMax H3 Director on fal now works with Clapper for AI filmmaking — flngr · 2026-09-06
- GPT-6 Astra first model to nail Apple's 107s Don't Blink video one-shot — socialwithaayan · 2026-09-06
- ComfyUI Workflow Puts Multi-Video Generation Results Into XYZ GridPlots — GeroldMeisinger · 2026-09-06