Principia Benchmark Exposes Poor Physical Reasoning in Video Models

The Principia benchmark tests whether video generation models truly understand Newtonian physics via calibration-free relational consistency between paired objects; mainstream models all scored below 0.42 on physical consistency, revealing a major reasoning gap.

2026-09-04 ~ 2026-09-06 · 3 related posts