Apple-PI benchmarks 11 video models on physics-grounded reasoning, and the best scores 0.473
_akhaliq · x · 2026-07-24
What the paper proposes
Apple-PI is a new benchmark for evaluating video models through the lens of physical law, instead of judging only whether the final output looks plausible.
Core components
- Orchard: a dataset of 400 videos across 10 classical mechanics tasks.
- Benchmark protocol: a three-stage reasoning flow — Perception, Formulation, Deduction — using chain-of-frames prompting on annotated first frames.
- Evaluation suite: combines subjective MLLM scoring with physics-law-grounded objective measures.
Main finding
Benchmarking 11 models shows current video models are still far from reliable law-grounded world simulators. The best model scores only 0.473. The authors argue failures cluster around a Perception → Formulation → Deduction bottleneck, weak multi-law transfer, and a persistent sim-to-real gap.
More from Multimodal
- Seedance 2.0 Test: Stunning Motion Transfer via Single Prompt — miilesus · 2026-07-24
- A local pipeline turns text or images into animated 3D assets in under 30 minutes — Tamerygo · 2026-07-24
- Microsoft Releases Mage-Flow-Edit-Turbo for Instruction-Based Image Editing — microsoft · 2026-07-24
- Yapper MCP brings vibe-video creation directly into Claude chat — Scobleizer · 2026-07-24
- Claude can split a 40-minute YouTube video into shorts and predict the winner — dr_cintas · 2026-07-24
- Creator outlines a Gemini Omni Flash workflow for realistic AI UGC videos — eptwts · 2026-07-24