ISMIR 2026 paper MPEcho adds phoneme conditioning to cut lyric errors in cover song generation

affige_yang · x · 2026-08-04

MPEcho improves lyric control in cover song generation with phoneme-level conditioning

The paper proposes MPEcho, a melody- and phoneme-aware framework for controllable cover song generation. It adds a phoneme encoder and length regulator to the SongEcho pipeline so the model can use explicit phoneme-level conditioning instead of relying only on voiced/unvoiced tags.

To support this, the authors also built Phonsa, a Whisper-based phoneme transcription pipeline for singing voice. According to the paper, Phonsa provides high-precision phoneme annotations and helps address the lack of paired audio-phoneme data. Experiments show that the approach substantially reduces phoneme error rate (PER) and improves end-to-end lyric controllability. Code, demo audio, and weights are available.

Related event: MPEcho Optimizes Cover Generation with Phoneme Timing(3 posts)→

Original post →

More from Multimodal

Multimodal channel →