Switch language한국어
Back to the list

Mistral AI just released a text-to-speech model it says beats ElevenLabs — and it's giving away the weights for free

TL;DR AI

Key summary

2 min read
  1. Mistral AI released the full model weights distributed Voxtral TTS weights free for download.

  2. Voxtral TTS built the backbone with 3.4 billion parameters transformer decoder backbone of 3.4 billion parameters.

  3. Said the model can run on laptops and smartphones quantized inference requires about 3 GB of RAM.

  4. Reported a time-to-first-audio of 90 milliseconds typical input yields 90 ms time-to-first-audio.

  5. Reported generation speed of about six times real-time speech generation at approximately 6x real-time speed.

Read the original