Principia benchmark: no video model exceeds 0.42 on physics despite VBench ~0.8
anand_bhattad · x · 2026-09-07
The Principia benchmark evaluates Newtonian physics in video models via calibration-independent relational consistency between paired objects, covering eight phenomena (gravity, friction, pendulum, etc.). Across thousands of generations from six state-of-the-art video generators, none exceeded 0.42 despite 0.8 VBench scores; the best VLM detected physics violations with only 67% accuracy.
Related event: Principia Benchmark Exposes Video Models' Poor Physics Understanding(3 posts)→
More from Multimodal
- CuePrecise: an open-source local MCP that turns long YouTube videos into searchable timelines — Forward-Associate523 · 2026-09-07
- Photorealistic character lip-sync gym video made with Flow AI — rajeev5059 · 2026-09-07
- "Astra, render this art in 3D and show me what's inside" — one prompt, done — blakesamic · 2026-09-07
- Midjourney prompt ported to Grok Imagine delivers surprisingly good results — michaelrabone · 2026-09-07
- 804-Page Manhua on One RTX 4080: Reverse-Engineering gpt-image's Whole-Page Generation — Sensitive-Wealth5801 · 2026-09-07
- One person made a 94-minute AI sci-fi film adapting Liu Cixin's 'Mountain' for ~$28k RMB — xiaosun86 · 2026-09-07