Apple-π: First Benchmark for Law-Grounded Physical Video Reasoning
jiqizhixin · x · 2026-08-10
While modern video generation models are hailed as world models with internalized physical laws, existing benchmarks only evaluate plausibility at the output level. The new Apple-π benchmark anchors evaluation explicitly in physical laws with three components:
- Orchard Dataset: 400 videos covering 10 classical mechanics tasks, separating single-law tasks for confounder-free diagnosis from multi-law tasks for generalization probing.
- 3-Stage Protocol: Based on scientific reasoning (Perception, Formulation, Deduction), treating generated video as the model's visible reasoning trace.
- Evaluation Suite: Combines MLLM-based subjective scoring with physics-law-grounded objective measures to pinpoint exactly where a model fails.
Benchmarking 11 models shows current video models are far from reliable law-grounded world simulators.
More from Research
- Theoretical Blind Spot of Discrete Diffusion: Fails to Learn Joint Probability Distributions — kalomaze · 2026-08-10
- NeurIPS 2026 Calls for Papers on Physical Understanding for Embodied AI — shaohua0116 · 2026-08-10
- ReASearch: Single LLM Agent Outperforms Specialized Optimizers Across ML Workflows — _reachsumit · 2026-08-10
- Sakana AI Summarizes 'AI Scientist' Progress in End-to-End Research Automation — SakanaAILabs · 2026-08-10
- NBER Paper Explores AI Agent Economics: Plunging Transaction Costs to Reshape Market Design — danielrock · 2026-08-10
- Meta Introduces SYF: An LLM-based Agentic System for Real-time Conversational Recommendations — _reachsumit · 2026-08-10