Google has released Gemini 3.5 Transcribe, a speech-to-text model designed for intelligent voice interactions that converts raw audio into accurate, polished text while handling background noise and disfluencies.
- Achieves an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use-cases as measured by Artificial Analysis.
- Improves time to final transcription by 70% compared to the previous Chirp 3 model.
- Supports over 85 languages with automatic detection and handles custom vocabulary for specialized jargon.
- Offers real-time streaming via the Live API and pre-recorded audio processing via the Interactions API.
- Enables function calling to delegate tasks like image generation to other Gemini models.
The model is available in public preview through the Gemini API, Google AI Studio, and the Gemini Enterprise Agent Platform.