Mistral AI just released a text-to-speech model it says beats ElevenLabs — and it's giving away the weights for free

TL;DR AI
2 min readKey summary
Mistral AI released the full model weights distributed Voxtral TTS weights free for download.
Voxtral TTS built the backbone with 3.4 billion parameters transformer decoder backbone of 3.4 billion parameters.
Said the model can run on laptops and smartphones quantized inference requires about 3 GB of RAM.
Reported a time-to-first-audio of 90 milliseconds typical input yields 90 ms time-to-first-audio.
Reported generation speed of about six times real-time speech generation at approximately 6x real-time speed.



