Google has released two new text-to-speech models, gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts, featuring a library of over 2,000 voices.
The API supports creating custom voices using just a 30-second audio sample and allows users to define multi-speaker conversations with distinct voice styles.
In testing, generating 1 minute and 18 seconds of audio took approximately 20 seconds and cost 2.74 cents using the standard Flash model.