ISMIR 2026 paper MPEcho adds phoneme conditioning to cut lyric errors in cover song generation
affige_yang · x · 2026-08-04
MPEcho improves lyric control in cover song generation with phoneme-level conditioning
The paper proposes MPEcho, a melody- and phoneme-aware framework for controllable cover song generation. It adds a phoneme encoder and length regulator to the SongEcho pipeline so the model can use explicit phoneme-level conditioning instead of relying only on voiced/unvoiced tags.
To support this, the authors also built Phonsa, a Whisper-based phoneme transcription pipeline for singing voice. According to the paper, Phonsa provides high-precision phoneme annotations and helps address the lack of paired audio-phoneme data. Experiments show that the approach substantially reduces phoneme error rate (PER) and improves end-to-end lyric controllability. Code, demo audio, and weights are available.
Related event: MPEcho Optimizes Cover Generation with Phoneme Timing(3 posts)→
More from Multimodal
- Open-sourced SARAS, an AI video platform that turns topics into full videos — sai_teja_ · 2026-08-04
- Video-analysis tool v0.5.1 adds direct AI analysis with Gemini, Kimi, OpenAI and Claude — sujingshen · 2026-08-04
- MM H3 local test on a 3090 shows 500–900 second renders and high heat — TensorTinkererTom · 2026-08-04
- MM H3 local test on a 3090 shows 500–900 second renders and high heat — TensorTinkererTom · 2026-08-04
- ChatGPT prompt workflow produces a 10-second Blender orbit animation at 768×768 and 24 fps — goodside · 2026-08-04
- MiniMax-H3 text-to-video GGUF weights trend on Hugging Face — realrebelai · 2026-08-04