WanSong: Pure Diffusion Model for 5-Minute Songs
Crazy-Repeat-2006 · reddit · 2026-08-15
WanSong v1.0 technical report released. It is a pure diffusion-based long-form music generation model capable of producing high-fidelity, dual-stem (vocals and accompaniment) songs up to 5 minutes in a single run. It simplifies cascaded pipelines and supports fast inference via step-distillation for commercial use.
More from Multimodal
- LTX 2.5 video generation quality shows stunning realism — tom_doerr · 2026-08-15
- Social media video concept for T-shirt brand using Minimax H3 r2v + Turbo LoRA — Time-Ad-7720 · 2026-08-15
- AI video quality enables niche 'The Office' episodes — venturetwins · 2026-08-15
- AI music video 'Wu Zetian' released: 36 shots, 3:24, English vocal version — Hongyi_AI · 2026-08-15
- Can ComfyUI generate 3D models locally? User seeks guide — jcam12312 · 2026-08-15
- MiniMax H3 video gen demo: 14 mins runtime, upscaled via RTX — FionaSherleen · 2026-08-15