New video-dynamics paper separates camera motion from object motion using frozen ViT features
FunAILab · hf · 2026-07-24
What the paper claims
The authors study how to recover structured motion representations from frozen image-ViT features, separating camera motion from object motion in video.
Key idea
They propose Structured Dynamics Model (SDM), which uses future-feature prediction to split dominant temporal change from residual dynamics instead of collapsing video motion into one entangled latent.
Training and evaluation
- Trained with self-supervised learning on real videos
- Adds weak supervision from synthetic Kubric data
- Introduces ProbeMotion, a new benchmark covering synthetic and real videos with camera motion, object motion, and combined dynamics
Results
SDM beats baselines built on global CLS or average-pooled features, and in several probes performs favorably even against strongly supervised representations such as VGGT, despite using much weaker supervision.
Why it matters
The work suggests pretrained image models can be repurposed into structured video-dynamics representations, giving video understanding systems a more useful inductive bias for motion analysis.
More from Research
- Cohere Labs opens a 48-hour model challenge on language learning and reasoning — Cohere_Labs · 2026-07-24
- SupraLabs releases a 5M-row reasoning corpus for tiny-model fine-tuning — LH-Tech_AI · 2026-07-24
- A factor-model paper shows panel causal inference without parallel trends — PtrPomorski · 2026-07-24
- Ordinary Least Squares regression, explained with a simple fit chart — mdancho84 · 2026-07-24
- Leaky Language Models show token timing can expose architecture and optimizations — chaumian · 2026-07-24
- Study finds OpenHandsDev used the least energy in a four-framework coding-agent test — rajistics · 2026-07-24