Switch language한국어
Back to the list

Closing the ‘Expressivity Gap’: How Mistral’s Voxtral TTS is Redefining Multilingual Voice Cloning with a Hybrid Autoregressive and Flow-Matching Architecture

TL;DR AI

Key summary

2 min read
  1. Mistral AI launched Voxtral TTS, its first text-to-speech model, as open weights and an API.

  2. The hybrid system combines an autoregressive decoder, a flow-matching acoustic model, and a custom codec to generate speech from short reference clips in nine languages.

  3. Mistral says Voxtral TTS improves speaker fidelity and expressive speech, with strong native-speaker evaluation results and low-latency serving on a single NVIDIA H200.

  4. The release targets production voice agents, audiobook narration, and multilingual use cases where current TTS systems often struggle with consistent, natural voices.

Read the original