Google has released Gemini 3.8 Flash TTS and Flash-Lite TTS via the Gemini API and AI Studio, achieving top rankings on Hume's Voice Design Benchmark and Voice Arena across six languages.
- Users can now replicate a voice by recording two consent-based clips or design a new one from a single sentence without prior recordings.
- The API supports inline sound effects like <short pause> and detailed delivery styles via speech_metadata.style.
- Replicated voices are reusable, while non-consented voices can be stored as encrypted keys expiring after seven days.
- Prompting has changed significantly: instructions inside text are spoken aloud, multi-speaker turns require explicit speaker tags, and long Audio Profile prompts may cause voice drift.
The update allows for precise control over synthetic voice generation, though existing prompts from the gemini-3.1-flash-tts-preview version will break without migration.