BUPT study: RoPE attention decay causes video diffusion models to violate physics
BUPT-CIST · hf · 2026-09-22
BUPT researchers present the first interpretability study of the "motion planning" process in text-to-video diffusion models. They find trajectories form in early denoising stages ("first shape, then details"), identify attention heads driving motion planning, and show Rotary Position Embedding (RoPE) induces excessive spatial attention decay that prematurely locks candidate regions into physically implausible positions. A lightweight fix—scaling RoPE frequencies across denoising steps—reduces the decay, and both training-free and training-based experiments confirm improved physical commonsense in generated videos.
More from Multimodal
- Pixel-art tour of London generated with MiniMax H3 shows consistent stylized video — Sufficient_Cause_43 · 2026-09-22
- Hunyuan Image 3.5 lands exclusively on OnSolo with 5 refs, 2K output, 1 credit — LearnWithBishal · 2026-09-22
- Reddit Asks for Real-World LongCat-Video Inference Times on RTX 4090 to H100 — Clean_Extreme_3970 · 2026-09-22
- Tencent ARC's WorldCrafter adds implicit 3D-aware memory to video world models — TencentARC · 2026-09-22
- Grok 4.7 made this in Blender — demo shows the model driving 3D software — iamfakhrealam · 2026-09-22
- Kyutai releases Voice of Reason, a speech-native reasoning model hitting 77.1% on GSM8K — alexcovo_eth · 2026-09-22