xAI has released Grok Voice Transcribe 2.0, a speech-to-text model that is twice as accurate as version 1.0 while maintaining the same pricing. Built on the audio foundation model behind Grok Voice, it leverages a unique dataset of live, noisy, multilingual audio to handle challenging real-world conditions like flaky phone lines and competing voices.
- Ranks first for accuracy among 32 streaming models on the public Artificial Analysis leaderboard.
- Improves word error rate on short phrases from 20.6% to 6.8%, marking its largest multilingual improvement.
- Leads every tested model on telephony audio and improves performance across four internal production datasets.
- Automatically detects language and handles mid-recording switches in a single pass for dozens of languages.
The update will soon become the default in the Speech-to-Text API, with version 1.0 scheduled for deprecation in the coming weeks.