m-a-p's YuE2 Unifies Symbolic and Audio Music Generation, Matching Suno v5
m-a-p · hf · 2026-09-29
m-a-p releases YuE2, unifying symbolic and audio music generation at frontier quality through symbolic planning.
- Architecture: a single AR-NAR Mixture-of-Transformers (MoT) writes a readable score (melody and harmony), expands it into semantic music tokens, and realizes full-song audio.
- Results: experts prefer symbolic planning 49.3% vs 34.6% overall in matched-checkpoint comparisons; YuE2 scores 6.73 on WildSongBench (SongBench Global Avg), beating all public baselines, and 6.96 with best-of-8. Expert listening finds it competitive with proprietary generators—best-of-8 beats Suno v4.5 and is nearly balanced against Suno v5.
- Supporting tech: MERT2 sets new SOTA on 14 of 15 MARBLE metrics; SheetSage2 leads 12 of 15 benchmark-metric pairs in lead-sheet transcription, enabling learning from recordings without aligned scores.
- Extras: score edits preserving unedited content, zero-shot covers, and agentic music editing via external LLMs translating user feedback into composition revisions.
More from Multimodal
- Early 'Opus 5.5' animation test drops: Into the ANTVERSE — ChrisGPT · 2026-09-29
- On DGX Spark, bf16 beats int8 convrot: H3 video gen 272s vs 287s in real tests — dtdisapointingresult · 2026-09-29
- One prompt, $20, 10 minutes: NoSpoon + Minimax H3 auto-shoots an AI cartoon episode — sachinmaya1980 · 2026-09-29
- Local I2V on a 4080 Super: Hours per 4-Second Clip or OOM — Thodane · 2026-09-29
- Reddit user fixes character drift in AI video with a retention_analysis prompt structure — SSj_Enforcer · 2026-09-29
- AMD 9070XT MiniMax image-to-video: ComfyUI tuning cuts generation from 58 to 21 minutes — Sofa-Sleuth · 2026-09-29