Wan-Dancer: Minute-Level Music-to-Dance Generation
pmttyji · reddit · 2026-07-13
Wan-Dancer: Minute-Level Music-to-Dance Generation
The authors introduce Wan-Dancer, a model designed to generate long, high-resolution, rhythm-synchronized dance videos from full music tracks. It aims to overcome the instability that current diffusion video models face around the 20-second mark.
Core Approach
- Decouples generation into two tiers: global keyframe planning + local temporal refinement to ensure long-range consistency.
- Utilizes the full music context to maintain structural coherence over long durations.
- Introduces time-mapped RoPE for dynamic frame rate adaptation, improving audio-visual alignment.
- Employs optical flow-based loss to enhance motion continuity.
- Uses motion-speed control to preserve details during fast movements.
Results & Availability
- The paper claims it outperforms traditional duration-limited methods in experiments.
- Capable of generating stable dance videos exceeding 1 minute at 720p/30fps.
- Covers 5 dance styles and supports both audio and text conditions.
- Model weights and inference code are open-sourced, with demos available on ModelScope Studio and Hugging Face Space.
Related event: Wan-Dancer-14B Open-Sourced for Music-Driven Long Dance Video Generation(6 posts)→
More from Multimodal
- Reddit user seeks ComfyUI NSFW text-to-image and image-to-video workflows under 20 GB VRAM — hobbyist2020 · 2026-07-22
- Krea 2 users recommend a two-pass Clownshark sampler setup for sharper image details — listopalafoto · 2026-07-22
- Gemini Omni Flash turns a boat cabin into a cave in Flow by Google — chrisfirst · 2026-07-22
- A simple workflow to turn a photo into an image prompt using Gemini, Grok, or GPT Image — harshitagu72595 · 2026-07-22
- A Reddit user proposes a consistency LoRA to keep anime and game scenes visually stable — ThirdWorldBoy21 · 2026-07-22
- Hand-painted figurines run through Seedance look eerily alive — cocktailpeanut · 2026-07-22