Switch language한국어
Back to the list

Google Releases Gemini 3.1 Flash Live: A Real-Time Multimodal Voice Model for Low-Latency Audio, Video, and Tool Use for AI Agents

TL;DR AI

Key summary

2 min read
  1. Gemini 3.1 Flash Live Google released Gemini 3.1 Flash Live in preview, released in preview for developers via the Gemini Live API in Google AI Studio.

  2. The model natively processes multimodal streams to reduce latency, designed for low-latency real-time voice interactions and native multimodal stream processing.

  3. Google positioned the model as its highest-quality audio and speech model to date, google described it as their highest-quality audio and speech model to date.

  4. Multimodal Live API the Multimodal Live API provides a stateful bi-directional streaming interface using WebSockets, supports stateful, bi-directional streaming over WebSockets (WSS).

  5. The API expects raw 16-bit PCM audio (16 kHz, little-endian) and returns raw PCM audio, accepts raw 16-bit PCM audio at 16 kHz little-endian as input and returns raw PCM audio output.

Read the original