Advancing voice intelligence with new models in the API

TL;DR AI
2 min readKey summary
OpenAI launched three new audio models in its API for voice applications.
The lineup includes GPT-Realtime-2 for reasoning-focused real-time voice, GPT-Realtime-Translate for live speech translation, and GPT-Realtime-Whisper for streaming speech-to-text.
These models give developers more tools to build voice agents that can listen, reason, translate, transcribe, and respond in real time.
The launch expands what’s possible for more natural, multilingual voice interfaces and live conversational products.



