Video diffusion models break physics because RoPE locks motion paths early, arXiv study finds

udmrzn · x · 2026-09-24

A new arXiv paper presents the first interpretability study of motion planning in text-to-video diffusion models, explaining why SOTA systems often violate physics.

Key findings:

Fix: a lightweight architectural modification that scales RoPE frequency across denoising steps. Both training-free and training-based experiments confirm improved physical consistency, without external priors or specialized data.

Related event: Study Explains Why Video Diffusion Models Break Physics(2 posts)→

Original post →

More from Multimodal

Multimodal channel →