WanSong Generates 5-Min Stem Songs
Justgototheeffinmoon · reddit · 2026-07-17
The Wan team released a technical report on arXiv detailing WanSong, a diffusion-based music generation model:
- Capable of generating complete songs up to 5 minutes long in a single pass.
- Simultaneously outputs both vocal and instrumental tracks during the same generation step.
- The authors emphasize this is a pure diffusion approach, not an autoregressive or multi-stage cascaded pipeline.
- The paper notes that inference can be further accelerated via step distillation.
- Supports fine-tuning / customization for downstream editing and remixing.
The post suggests that if the claims of "stem output + fine-tuning" hold up to independent verification, this type of foundational model will greatly benefit teams building music editors, plugins, and workflow tools, as it is much easier to integrate than black-box mixed audio outputs.
Related event: Wan Team Releases WanSong Music Generation Model(2 posts)→
More from Multimodal
- Midjourney prompt turns a bee into a glitching pixel explosion — michaelrabone · 2026-07-22
- A physics reward can improve video generation without creating a real physics engine — Dapper-Drawer4546 · 2026-07-22
- HeyGen adds a media-sourcing skill for coding agents with 75k images and 10k tracks — HeyGen · 2026-07-22
- Reddit user seeks ComfyUI NSFW text-to-image and image-to-video workflows under 20 GB VRAM — hobbyist2020 · 2026-07-22
- Krea 2 users recommend a two-pass Clownshark sampler setup for sharper image details — listopalafoto · 2026-07-22
- Gemini Omni Flash turns a boat cabin into a cave in Flow by Google — chrisfirst · 2026-07-22