WanSong Generates 5-Min Stem Songs
Justgototheeffinmoon · reddit · 2026-07-17
The Wan team released a technical report on arXiv detailing WanSong, a diffusion-based music generation model:
- Capable of generating complete songs up to 5 minutes long in a single pass.
- Simultaneously outputs both vocal and instrumental tracks during the same generation step.
- The authors emphasize this is a pure diffusion approach, not an autoregressive or multi-stage cascaded pipeline.
- The paper notes that inference can be further accelerated via step distillation.
- Supports fine-tuning / customization for downstream editing and remixing.
The post suggests that if the claims of "stem output + fine-tuning" hold up to independent verification, this type of foundational model will greatly benefit teams building music editors, plugins, and workflow tools, as it is much easier to integrate than black-box mixed audio outputs.
Related event: Wan Team Releases WanSong Music Generation Model(2 posts)→
More from Multimodal
- Reddit user chains Ideogram 4 and Krea2 to mimic bbox-based image positioning — v3lh0t05c0 · 2026-07-22
- Ultimate Face Fix: Open-Source Multi-Face Repair Node for ComfyUI — Merserk13 · 2026-07-22
- Getting Started with AI Video: Solving Consistency and Censorship — cynicalnewenglander · 2026-07-22
- Storyboard-first workflows are making AI dance videos and influencers more consistent — aftahi_ai · 2026-07-22
- Runpod MCP and Claude help spin up image and video generation workflows — 802high · 2026-07-22
- Midjourney prompt turns a bee into a glitching pixel explosion — michaelrabone · 2026-07-22