Switch language한국어
Back to the list

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue

TL;DR AI

Key summary

2 min read
  1. Researchers introduced SwanData-Speech and SwanVoice, a zero-shot TTS pipeline for 1–4 speakers.

  2. It combines pause-aware text conditioning, a VAE, a flow-matching DiT, and diffusion-based post-training to improve expressive long-form speech.

  3. On SwanBench-Speech, it beat open-source baselines in richness and hierarchy, especially for monologue and dialogue generation.

  4. The system still has room to improve on content accuracy, but it advances coherent, emotionally consistent multi-speaker speech synthesis.

Read the original