AudioSlopServer: host multiple audio diffusion models on one GPU

SteveLittleFish · reddit · 2026-09-20

A Redditor open-sourced AudioSlopServer (permissive license), a service that hosts multiple audio generation models on a single GPU.

How it works:

Included: YuE 2, ACE-Step 1.5 XL, Stable Audio 3, DEMUCS (stem separation), and a forced aligner for lyrics. TTS and STT are planned toward a full AI radio station.

Example workflow: generate a song with ACE-Step → separate vocals/backing → reprocess the backing through Stable Audio 3 → mix vocals back in, fixing ACE-Step's odd guitar rendering.

Original post →

More from Multimodal

Multimodal channel →