Apple-PI benchmarks 11 video models on physics-grounded reasoning, and the best scores 0.473

_akhaliq · x · 2026-07-24

What the paper proposes

Apple-PI is a new benchmark for evaluating video models through the lens of physical law, instead of judging only whether the final output looks plausible.

Core components

Main finding

Benchmarking 11 models shows current video models are still far from reliable law-grounded world simulators. The best model scores only 0.473. The authors argue failures cluster around a Perception → Formulation → Deduction bottleneck, weak multi-law transfer, and a persistent sim-to-real gap.

Original post →

More from Multimodal

Multimodal channel →