MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation

TL;DR AI
2 min readKey summary
Researchers introduced MPEcho, a cover song generation framework that adds phoneme-level conditioning and timing control to SongEcho.
MPEcho includes a phoneme encoder and length regulator to better preserve lyrics, pronunciation, and melody in synthesized singing.
The team also developed Phonsa, a Whisper-based transcription model that produces phoneme-level singing annotations.
Together, these tools improve controllable cover song generation and help reduce lyric and pronunciation errors.
