WanSong: 5-Minute Long Song Generation Model
Wan-AI · hf · 2026-07-17
## WanSong v1.0: A Pure Diffusion Model for Long Song Generation Wan-AI released the technical report for **WanSong v1.0**, aiming for "commercial-grade" long-form song generation. The authors believe the current difficulty in music generation lies in achieving high fidelity, long duration, and controllability/customization simultaneously. ### Model Features - Adopts a **pure diffusion** approach, rather than AR or cascaded multi-stage pipelines - Can directly generate high-fidelity, multilingual songs up to **5 minutes** long - A single generation outputs dual tracks: **vocals + accompaniment** - Supports inference speedup via **step-distillation** - Provides highly efficient fine-tuning and customization paths for downstream editing tasks This report focuses not on isolated demos, but on pushing the efficiency, quality, and controllability of long-form music generation forward together.
Related event: Wan Team Releases WanSong Music Generation Model(2 posts)→
More from Multimodal
- OCT-Bench sets 10,076 questions to test whether multimodal models really understand retinal scans — Baochen Fu · 2026-07-21
- LTX-2.3 face-and-voice LoRA training can work on 12GB VRAM with heavy tradeoffs — __alpha_____ · 2026-07-21
- Seedance 2.0 turns one reference image into a cinematic fight scene — techhalla · 2026-07-21
- Seedance 2.0 keeps character consistency across 15+ shots with just 3 prompts — techhalla · 2026-07-21
- DecartAI’s Lucy 2.5 Realtime lands on fal with live video-to-video editing — gorkem · 2026-07-21
- Google Gemini’s Omni text-to-video output is getting better, user says — michaelrabone · 2026-07-21