Vector's 3DSPA reward signal fixes physics violations in generative video models
VectorInst · x · 2026-10-08
At Vector Institute's 2026 AI Summit, researcher Kelsey Allen showed how cognitive science exposes gaps in modern AI.
- LLMs excel at knowledge representation but lack physical reasoning: the Virtual Tools game shows humans solve novel physics problems in a few attempts via accurate mental simulation — a capability current models lack.
- The team proposed 3DSPA (3D Semantic Point Autoencoder), which predicts approximate 3D point trajectories to measure physical realism in generated video without ground-truth comparisons, detecting physics violations like a human infant would.
- Used as an RL reward signal, 3DSPA measurably improves physical realism: gravity behaves correctly and objects stop passing through solid barriers.
- The talk also raised how improving AI reshapes human cognition and fosters overreliance; full talk on Vector's YouTube channel.
More from Multimodal
- ExploreNet ablation: perturbing learned high-sensitivity channels drives larger human-perceived change — StellaLisy · 2026-10-08
- ExploreNet outperforms FlowGRPO by learning state-dependent exploration noise — StellaLisy · 2026-10-08
- Fish Audio's Drama 3 lets you direct voice acting in plain language, shifts emotion mid-line — omarsar0 · 2026-10-08
- Single-Word Midjourney Prompt: 'Threnody' Yields Mournful AI Art (Series 3/7) — tisch_eins · 2026-10-08
- 500+ Opus 5.5 generated videos with copy-paste prompts, effort Max recommended — yihui_indie · 2026-10-08
- Video world models flunk physics: 8 SoTA models top out at 57.76/100 on new benchmark — JaynitMakwana · 2026-10-08