NTU & CUHK Release Apple-π to Test if Video Models Truly Understand Physics

机器之心 · wechat · 2026-07-30

A joint team from NTU and CUHK released Apple-π, a benchmark designed to evaluate whether video generation models genuinely understand physical laws rather than merely mimicking visual statistical regularities.

The benchmark includes 400 classical mechanics cases and breaks down scientific reasoning into three stages: perception, law formulation, and deduction. Tests on 11 models reveal a clear "reasoning funnel" where perception is easiest and deduction is hardest. The best video model scored only 0.473, and even top unified models struggled with dynamic deduction, highlighting a significant gap before video models can become reliable world simulators.

Original post →

More from Multimodal

Multimodal channel →