Google has released two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, as part of its Gemini Audio family. These models introduce generative voice design from natural language prompts and line-by-line performance direction via the Gemini API and Google AI Studio.
- Gemini 3.8 Flash TTS targets creative direction for gaming and immersive media, ranking #1 on the Hume AI Voice Design Benchmark with a score of 71.4.
- Gemini 3.8 Flash-Lite TTS is optimized for high-volume, cost-efficient production like dubbing and ranks #2 on the Hume AI Overall Quality Index.
- The release expands the voice library to over 2,000 production-ready voices and enables generative design across more than 100 languages.
- Both models support native two-speaker staging, vocal bursts, backchanneling, and long-form generation with consistent quality.
- Voice replication requires a 30-second sample and verbal consent, with all outputs watermarked by SynthID.
The models provide developers with granular control over tone, pacing, and character design while offering scalable options for different production needs.