New video-dynamics paper separates camera motion from object motion using frozen ViT features

FunAILab · hf · 2026-07-24

What the paper claims

The authors study how to recover structured motion representations from frozen image-ViT features, separating camera motion from object motion in video.

Key idea

They propose Structured Dynamics Model (SDM), which uses future-feature prediction to split dominant temporal change from residual dynamics instead of collapsing video motion into one entangled latent.

Training and evaluation

Results

SDM beats baselines built on global CLS or average-pooled features, and in several probes performs favorably even against strongly supervised representations such as VGGT, despite using much weaker supervision.

Why it matters

The work suggests pretrained image models can be repurposed into structured video-dynamics representations, giving video understanding systems a more useful inductive bias for motion analysis.

Original post →

More from Research

Research channel →