CosyEdit2: Speech-Editing-Oriented Reinforcement Learning Unlocks Better Zero-Shot TTS

TL;DR AI
2 min readKey summary
Researchers introduced CosyEdit2, a two-stage post-training method for speech editing.
It first uses supervised editing training, then applies editing-oriented GRPO on target-speech-free data.
The approach improves speech editing performance and also boosts zero-shot TTS quality.
The results suggest speech editing optimization and zero-shot TTS are more closely linked than previously thought.
