YuE2-3B: Open Music Model With Editable Scores Beats Suno v5 on WildSongBench
realmrfakename · x · 2026-09-10
Open music generation model YuE2-3B rivals Suno v5: turn lyrics and a style prompt into a complete song with vocals and accompaniment, then shape melody and chords through an editable score. On WildSongBench, YuE2 (best-of-8) achieves the highest SongBench average among all evaluated open and proprietary models: 6.9632 vs 6.8721 for Suno v5.
Key features:
- Compose and edit: melody + chords, melody-only, or direct generation; bring your own ABC score
- Agent-based editing: turn musical feedback into score, style, and lyric revisions, then render the next version
- Run locally: 48 kHz stereo songs on a 24GB GPU without quantization
- Architecture: an AR–NAR Mixture-of-Transformers backbone writes the score and semantic tokens, then generates acoustic latents via flow matching; a VAE converts them to stereo audio
Licensed CC-BY-NC-4.0 (non-commercial), with Hugging Face loading, CFG text guidance, and separate planning/synthesis APIs.
Related event: Open-Source Music Model YuE2 Released with Editable Sheet Music(3 posts)→
More from Multimodal
- Programmable World Model Separates Explicit State Evolution From Video Generation — Zheng-Hui Huang · 2026-09-10
- ByteDance's AgenticGen Uses Business Feedback and Human Rewards to Guide Ad Video Generation — ByteDance · 2026-09-10
- Flash-BoN: cheap drafts beat guided search for diffusion inference-time scaling, +8% AUC at scale — gowthami_s · 2026-09-10
- Alibaba open-sources 2 more speech enhancement models, completing a 7-model toolkit with 22ms echo cancellation — aigclink · 2026-09-10
- A 78-second AI food comedy: Pig Girl crashes a Viking winter feast — Substantial_Depth744 · 2026-09-10
- Shareable ChatGPT Prompt Turns Photos into Mecha-Style Art — Promptmethus · 2026-09-10