Can Local Music Model Yue2 Do True Reference2Audio? Community Asks
Ambitious_Fold_2874 · reddit · 2026-09-18
A Reddit user asks whether Yue2 or any local music model can do true reference2audio — generating audio conditioned on a reference clip's context plus user input, similar to MiniMax's ref2va capability.
The poster observes that Yue2 appears to only extract notes and melody from the reference audio and then generate music from that, rather than doing genuinely context-aware generation conditioned on the reference. The thread discusses whether local alternatives exist.
More from Multimodal
- Pika simplifies AI creation: pick what you want to make, not which model to use — minchoi · 2026-09-18
- Another Qwen image model appears imminent, hints Andrew Carr — andrew_n_carr · 2026-09-18
- Video DeltaNet hybrid attention speeds livestream video generation 14.5x on 8x B200 — Haocheng Xi · 2026-09-18
- UFO: unified omni-condition alignment evaluation for multimodal image generation, +15.25% human correlation — ustc-community · 2026-09-18
- AI-Generated Ming Dynasty Dynasty Reenactment Wows Viewers Online — angadc · 2026-09-18
- Creator builds StyleGAN dataset from ~3,000 synthetic JAX diffusion images — makeitrad1 · 2026-09-18