Apple-π Benchmark Reveals Shortcomings in Video Models' Physical Reasoning

Apple-π (Apple-PI) is a newly proposed benchmark designed to evaluate whether video generation models can reason based on explicit physical laws rather than merely assessing visual realism. The benchmark requires models to undergo a process from perception to formulation and deduction, serving as an auditable way to test their "physical intelligence." Test results indicate that current video models still face significant limitations in acting as world simulators.

Confirmed

The benchmark includes a dataset named Orchard, consisting of 400 videos covering 10 different physical concepts. After evaluating 11 video models, it was found that the best-performing model achieved a physical reasoning accuracy of only 0.473. Currently, the paper, project page, and code for the project have all been made public.

Why it matters

There is extensive industry discussion regarding the potential of video generation models to act as "world simulators." Apple-π provides a standardized diagnostic tool, revealing that even with highly realistic video generation, models still exhibit clear shortcomings in underlying physical logic reasoning. This points the way for future improvements in the physical consistency of video models.

2026-07-24 ~ 2026-07-25 · 5 related posts

Primary sources

2 near-duplicate retellings: liuziwei7 · _akhaliq