Video diffusion models break physics because RoPE locks motion paths early, arXiv study finds
udmrzn · x · 2026-09-24
A new arXiv paper presents the first interpretability study of motion planning in text-to-video diffusion models, explaining why SOTA systems often violate physics.
Key findings:
- Motion follows a 'first shape, then details' pattern: trajectories are committed during early denoising steps.
- A specific subset of attention heads, identified via cross-attention trajectory patterns and causal head contributions, drives motion planning.
- Rotary Position Embedding (RoPE) causes excessive spatial attention decay, prematurely locking candidate regions into physically implausible positions and suppressing coherent motion in adjacent frames.
Fix: a lightweight architectural modification that scales RoPE frequency across denoising steps. Both training-free and training-based experiments confirm improved physical consistency, without external priors or specialized data.
Related event: Study Explains Why Video Diffusion Models Break Physics(2 posts)→
More from Multimodal
- Hands-on with Pexo: an AI video agent that iterates your launch video via chat — HeyZoyaKhan · 2026-09-24
- Fashion video made locally with MiniMax H3: h3.c cuts Apple Silicon gen time 2-3x — MindfulPornographer · 2026-09-24
- Runway integrates its models directly into DaVinci Resolve timelines — runwayml · 2026-09-24
- Swapping Characters into a Hotel Lobby Video with Mitte's Low Mode — mso96 · 2026-09-24
- An AI-Created 'Lipstick Joke' Short Video — Klitolovac · 2026-09-24
- Building AI video workflows with Codex driving Dreamina CLI: node-level iteration without regenerating — HeyNayeem · 2026-09-24