ComfyUI Workflow: Generating Instrumental Music with Stable Audio 3 & ACE-Step
wjc_5 · reddit · 2026-08-02
The author shares a ComfyUI workflow that transforms short musical ideas and target durations into prompts, using both Stable Audio 3 and ACE-Step 1.5 XL to generate two instrumental tracks from the same plan.
Model Comparison:
- Stable Audio 3: More flexible across various styles, handling retro 8-bit game music and cyberpunk BGM naturally. However, longer tracks tend to expose noticeable repetition issues.
- ACE-Step 1.5 XL: Better for arrangements requiring clear structural changes. Its instrumental structure script can separately describe the intro, theme, variation, build, climax, and outro.
Technical Details:
The workflow uses a local text-generation node to create a strict JSON music plan, which is then split into specific prompts, BPM, time signature, and key. If local generation is too slow, users can bypass it by using a web LLM and pasting the JSON back. The author recommends starting with 2–3 minute tracks and has uploaded a comprehensive YouTube tutorial.
More from Multimodal
- Seeking Audio Upscaling LLMs: Is There a 'Super-Resolution' Model for Music? — LeatherRub7248 · 2026-08-24
- Describe your dream world to an AI dragon, which generates the planet for you — repligate · 2026-08-24
- Using kintsugi texture to fix cracks in edited 3D meshes — repligate · 2026-08-24
- Generating Hannibal Character Videos with FL2VA Model — Nimblecloud13 · 2026-08-24
- MiniMax H3 Revives Medieval Short Stories: Complete Workflow Shared — zanatas · 2026-08-24
- NAPE Audio Pretraining Achieves SOTA Without Decoders — kastnerkyle · 2026-08-24