Apple-PI benchmarks 11 video models on physics-grounded reasoning and finds a 0.473 ceiling
mmlab-ntu · hf · 2026-07-21
Apple-PI introduces a physics-grounded benchmark for video models
The paper argues that current video benchmarks mostly judge whether outputs look physically plausible, not whether the model arrives there through faithful reasoning. To address that gap, the authors introduce Apple-PI, a benchmark explicitly anchored in physical laws.
What it contains
- Orchard: 400 videos across 10 canonical classical-mechanics tasks.
- A three-stage protocol: Perception → Formulation → Deduction.
- Chain-of-frames prompting on infographic-annotated first frames, treating generated video as a visible reasoning trace.
- A hybrid evaluation suite combining MLLM-based subjective scoring with physics-law-grounded objective measures.
Main findings
- Evaluating 11 models, the best video model scores only 0.473.
- The authors find a bottleneck from perception to formulation to deduction.
- Multi-law state transfer remains weak.
- A persistent sim-to-real gap suggests current video models are still far from reliable law-grounded world simulators.
Related event: Apple-π Benchmark Tests Video Models' Physics Reasoning(3 posts)→
More from Multimodal
- Reddit user chains Ideogram 4 and Krea2 to mimic bbox-based image positioning — v3lh0t05c0 · 2026-07-22
- Ultimate Face Fix: Open-Source Multi-Face Repair Node for ComfyUI — Merserk13 · 2026-07-22
- Getting Started with AI Video: Solving Consistency and Censorship — cynicalnewenglander · 2026-07-22
- Storyboard-first workflows are making AI dance videos and influencers more consistent — aftahi_ai · 2026-07-22
- Runpod MCP and Claude help spin up image and video generation workflows — 802high · 2026-07-22
- Midjourney prompt turns a bee into a glitching pixel explosion — michaelrabone · 2026-07-22