WanSong Generates 5-Min Stem Songs

Justgototheeffinmoon · reddit · 2026-07-17

The Wan team released a technical report on arXiv detailing WanSong, a diffusion-based music generation model:

The post suggests that if the claims of "stem output + fine-tuning" hold up to independent verification, this type of foundational model will greatly benefit teams building music editors, plugins, and workflow tools, as it is much easier to integrate than black-box mixed audio outputs.

Related event: Wan Team Releases WanSong Music Generation Model(2 posts)→

Original post →

More from Multimodal

Multimodal channel →