AudioSlopServer: host multiple audio diffusion models on one GPU
SteveLittleFish · reddit · 2026-09-20
A Redditor open-sourced AudioSlopServer (permissive license), a service that hosts multiple audio generation models on a single GPU.
How it works:
- Forks of official repos with simple APIs added
- RAM offloading so model weights can move to system memory
- An orchestrator hot-swaps models; those that can't offload simply stop
- Each sub-service is its own Docker image, easing upstream merges
- Includes a test UI per service
Included: YuE 2, ACE-Step 1.5 XL, Stable Audio 3, DEMUCS (stem separation), and a forced aligner for lyrics. TTS and STT are planned toward a full AI radio station.
Example workflow: generate a song with ACE-Step → separate vocals/backing → reprocess the backing through Stable Audio 3 → mix vocals back in, fixing ACE-Step's odd guitar rendering.
More from Multimodal
- Indie dev crafts cursed 1980s VHS commercial with Midjourney and Kling AI to promote his game — sevencavesinteractiv · 2026-09-20
- MiniMax H3 3-step LoRA tested: ~13 min vs 48 min for Fused Turbo, with a small quality tradeoff — cgpixel23 · 2026-09-20
- Reskinning a 2-hour film with MiniMax H3 Max now costs roughly $1,500-$3,000 — bennash · 2026-09-20
- Qwen Image 2.1 hands-on: strong reference consistency, transparent background generation, minor drift — 13baaphumain · 2026-09-20
- AI video tool's 'next episode' button auto-generates endless bingeable episodes — Kyrannio · 2026-09-20
- NoSpoon microdrama agent auto-generates posters via Grok Imagine — Kyrannio · 2026-09-20