Switch language한국어
Back to the list

OpenAI unveils new voice model “GPT-Realtime-2,” with instant translation and low-latency transcription

TL;DR AI

Key summary

2 min read
  1. OpenAI introduced three new Realtime API speech models: GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper.

  2. The models aim to improve voice reasoning, live multilingual translation, and streaming speech recognition with lower latency and larger context windows.

  3. OpenAI says the new release outperforms GPT-Realtime-1.5 on benchmarks such as Big Bench Audio and Audio MultiChallenge.

  4. The update broadens OpenAI’s developer tools for building more natural voice products for support, education, media, and other live communication use cases.

Read the original