Principia: video models know less high-school physics than you'd think
mariyaivasileva · x · 2026-09-25
- The Principia benchmark (accepted at NeurIPS 2026 with three 5s, but no Spotlight/Oral) tests how well video generation models grasp physics.
- Core question: "What do video models know about high school physics? Less than you'd think." It covers 8 physical laws across 500+ real, carefully calibrated scenes, using relative evaluation with objective metrics.
- Building the dataset required enormous patience: calibrated experiments, verifying physical properties, even machining parts for comparability.
- The author draws a parallel to CVPR 2024's overlooked paper "Shadows Don't Lie and Lines Can't Bend".
More from Models
- Theo: $200 Claude Code plan now clearly beats Codex, weeks after trailing it — ssh4net · 2026-09-25
- Reddit asks: what's the real advantage of routing smaller model jev into LLM prompts — sogo00 · 2026-09-25
- Could the Jev model add reasoning offloading and N-gram augmentation? Fan speculation — ProposalOrganic1043 · 2026-09-25
- Redditor: models are now good enough that I stopped caring about prompt engineering — Large-Excitement6573 · 2026-09-25
- User says OpenAI's Sol 6 wastes their time, switches to $100 Claude plan and prefers it — AirportEither2456 · 2026-09-25
- Nagi-ENORMOUS beats Jev, Semif, Laya on Game Arena Benchmark — No_Skill_8393 · 2026-09-25