Google has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, native speech-to-speech models designed for real-time voice agents that reason and execute tools without breaking conversational flow.

  • Both models are now live in the Gemini Live API and Google AI Studio as hosted solutions.
  • Gemini 3.8 Live focuses on scale and cost efficiency, while Extended Thinking adds multi-step reasoning capabilities.
  • Extended Thinking ranks #1 on Artificial Analysis’ Speech to Speech Quality Index with a score of 82.6.
  • The models support asynchronous function calling, visual context, alphanumeric precision, and 97 languages.
  • Pricing is set at $0.005/min for audio input and $0.018/min for audio output via the Live API.

These releases provide a streamlined alternative to cascaded speech pipelines, enabling developers to build production-grade voice agents with integrated tool use and low latency.