Google has released two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, expanding the Gemini Audio family to offer dynamic voice generation capabilities for creators, developers, and enterprises.
- Gemini 3.8 Flash TTS enables deep creative direction, allowing users to create new voices from scratch using natural language prompts and control performance cues line by line.
- Gemini 3.8 Flash-Lite TTS is optimized for high-volume, cost-efficient scale, supporting dubbing and expressive voice agents with fine-grained control over tone and pacing.
- The models support over 100 languages and dialects, including regional varieties like Mexican Spanish and Scots English, and offer a library of 2,000+ production-ready voices.
- Voice replication features allow users to recreate vocal profiles from a 30-second sample, backed by consent verification, SynthID watermarking, and C2PA credentials.
- Both models secured top positions on Hume AI’s benchmarks, with Gemini 3.8 Flash TTS leading the Overall Quality Index and Voice Design Benchmark.
These tools are now available in Google AI Studio and the Gemini API for developers, while enterprise access is coming soon via Gemini Enterprise.