Can Local Music Model Yue2 Do True Reference2Audio? Community Asks

Ambitious_Fold_2874 · reddit · 2026-09-18

A Reddit user asks whether Yue2 or any local music model can do true reference2audio — generating audio conditioned on a reference clip's context plus user input, similar to MiniMax's ref2va capability.

The poster observes that Yue2 appears to only extract notes and melody from the reference audio and then generate music from that, rather than doing genuinely context-aware generation conditioned on the reference. The thread discusses whether local alternatives exist.

Original post →

More from Multimodal

Multimodal channel →