Apple-PI benchmarks 11 video models on physics-grounded reasoning and finds a 0.473 ceiling
mmlab-ntu · hf · 2026-07-21
Apple-PI introduces a physics-grounded benchmark for video models
The paper argues that current video benchmarks mostly judge whether outputs look physically plausible, not whether the model arrives there through faithful reasoning. To address that gap, the authors introduce Apple-PI, a benchmark explicitly anchored in physical laws.
What it contains
- Orchard: 400 videos across 10 canonical classical-mechanics tasks.
- A three-stage protocol: Perception → Formulation → Deduction.
- Chain-of-frames prompting on infographic-annotated first frames, treating generated video as a visible reasoning trace.
- A hybrid evaluation suite combining MLLM-based subjective scoring with physics-law-grounded objective measures.
Main findings
- Evaluating 11 models, the best video model scores only 0.473.
- The authors find a bottleneck from perception to formulation to deduction.
- Multi-law state transfer remains weak.
- A persistent sim-to-real gap suggests current video models are still far from reliable law-grounded world simulators.
Related event: Apple-π Benchmark Tests Video Models' Physics Reasoning(3 posts)→
More from Multimodal
- Non-coder builds full-featured Android ComfyUI client with ChatGPT, submits to Google Play — ComfierUI · 2026-09-11
- FastH3-Live hits 22fps: acceleration node benchmarks and the --vram-headroom trick — spartong945 · 2026-09-11
- Midjourney style code share: --sref 2912175708 — tisch_eins · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- MiniMax H3 MAX nails cooking anime clips: 15-second curry demo with prompts shared — Hailuo_AI · 2026-09-11
- MiniMax Music Production Toolkit 2.5 for ComfyUI adds full mastering chain — Vivid_Promise1700 · 2026-09-11