Closing the ‘Expressivity Gap’: How Mistral’s Voxtral TTS is Redefining Multilingual Voice Cloning with a Hybrid Autoregressive and Flow-Matching Architecture

TL;DR AI
2 min readKey summary
Mistral AI launched Voxtral TTS, its first text-to-speech model, as open weights and an API.
The hybrid system combines an autoregressive decoder, a flow-matching acoustic model, and a custom codec to generate speech from short reference clips in nine languages.
Mistral says Voxtral TTS improves speaker fidelity and expressive speech, with strong native-speaker evaluation results and low-latency serving on a single NVIDIA H200.
The release targets production voice agents, audiobook narration, and multilingual use cases where current TTS systems often struggle with consistent, natural voices.
