Apple-π Benchmark Tests Video Models' Physics Reasoning
Researchers introduced Apple-π, a benchmark of 400 videos testing if video models truly understand physics. The evaluation reveals that current models struggle with physical reasoning, hitting a maximum accuracy of only 0.473.
2026-07-21 ~ 2026-07-21 · 3 related posts
- Apple-PI benchmarks 11 video models on physics-grounded reasoning and finds a 0.473 ceiling — mmlab-ntu · 2026-07-21
- Apple-π benchmark asks whether video models reason about physical laws or just mimic motion — liuziwei7 · 2026-07-21
1 near-duplicate retellings: liuziwei7