ComfyUI-Olm-YuE2: Staged YuE2 Music Generation With an Editable Score Workflow
imlo2 · reddit · 2026-09-11
A developer released ComfyUI-Olm-YuE2, integrating the newly released open music model YuE2 into ComfyUI with far more than a single Generate button:
- Staged pipeline: generation split into Plan > Semantic > Synthesize > Decode, with intermediate results inspectable, savable, reusable and branchable.
- Optional score UI: sidebar rendered notation, raw ABC score view, an ABC editor node, and the ability to generate a plan, edit the score, then continue generation from the edited version.
- Input/output: style + lyrics in, 48 kHz stereo songs out.
- Engineering details: reuses ComfyUI's Torch/Transformers, needing only tiktoken and accelerate on top; includes measured VRAM numbers (tested on RTX 5090) and an experimental offload; weights not bundled. Environment: CUDA 13.0, Python 3.13.
Open source on GitHub; the author is seeking feedback from 16/24 GB GPU users.
More from Multimodal
- Beating Img2Video ghosting: 3 chained 4-second clips beat one 12-second render, twice as fast — Key_Education4018 · 2026-09-11
- Face Capture System v3 Shows Signs of Life — andrew_n_carr · 2026-09-11
- Single-Word Midjourney Prompt Series Uses Scots Word 'Smeddum' for Striking Imagery — tisch_eins · 2026-09-11
- One-shot animation with GPT-6 Astra wows users as model demos keep impressing — paw_lean · 2026-09-11
- The prompt behind that 75M-view AI volcano video is now out — charis_ai · 2026-09-11
- GPT-6 Astra + Magnific MCP: directing motion-graphics videos via conversation — charis_ai · 2026-09-11