BUPT study: RoPE attention decay causes video diffusion models to violate physics

BUPT-CIST · hf · 2026-09-22

BUPT researchers present the first interpretability study of the "motion planning" process in text-to-video diffusion models. They find trajectories form in early denoising stages ("first shape, then details"), identify attention heads driving motion planning, and show Rotary Position Embedding (RoPE) induces excessive spatial attention decay that prematurely locks candidate regions into physically implausible positions. A lightweight fix—scaling RoPE frequencies across denoising steps—reduces the decay, and both training-free and training-based experiments confirm improved physical commonsense in generated videos.

Original post →

More from Multimodal

Multimodal channel →