Introducing next-generation audio models in the API

TL;DR AI
2 min readKey summary
OpenAI has released next-generation audio models in its API.
The new gpt-4o-transcribe and gpt-4o-mini-transcribe models improve speech recognition accuracy and language detection.
The text-to-speech side also adds more natural, customizable speaking styles.
OpenAI says the models outperform Whisper on benchmarks and are built for more reliable voice applications.



