WanSong Generates 5-Min Stem Songs
Justgototheeffinmoon · reddit · 2026-07-17
The Wan team released a technical report on arXiv detailing WanSong, a diffusion-based music generation model:
- Capable of generating complete songs up to 5 minutes long in a single pass.
- Simultaneously outputs both vocal and instrumental tracks during the same generation step.
- The authors emphasize this is a pure diffusion approach, not an autoregressive or multi-stage cascaded pipeline.
- The paper notes that inference can be further accelerated via step distillation.
- Supports fine-tuning / customization for downstream editing and remixing.
The post suggests that if the claims of "stem output + fine-tuning" hold up to independent verification, this type of foundational model will greatly benefit teams building music editors, plugins, and workflow tools, as it is much easier to integrate than black-box mixed audio outputs.
Related event: Wan Team Releases WanSong Music Generation Model(2 posts)→
More from Multimodal
- FLUX.2 Klein Drifts Hard on Character Expressions While Free Gemini Holds Likeness — wacomlover · 2026-09-11
- Tencent Hunyuan releases AuK code and weights on GitHub with ComfyUI and fine-tuning support — aigclink · 2026-09-11
- Creator turns Bahamut vs Tiamat rivalry into an AI cinematic battle with Midjourney, GPT Image 2 and Seedance — azed_ai · 2026-09-11
- invideo launches AI agent-powered editor to automate repetitive editing tasks — azed_ai · 2026-09-11
- fable 5.1 recreates The Starry Night with 256,157 JavaScript brush strokes — cedric_chee · 2026-09-11
- GPT-6 Astra + Hyper3D Rodin MCP Generates 3D Assets in One Agent Flow — ahuja_priyank · 2026-09-11