xAI has released Grok Voice Transcribe 2.0, a speech-to-text model that is twice as accurate as version 1.0 while maintaining the same pricing. Built on the audio foundation model behind Grok Voice, it leverages a unique dataset of live, noisy, multilingual audio to handle challenging real-world conditions like flaky phone lines and competing voices.

  • Ranks first for accuracy among 32 streaming models on the public Artificial Analysis leaderboard.
  • Improves word error rate on short phrases from 20.6% to 6.8%, marking its largest multilingual improvement.
  • Leads every tested model on telephony audio and improves performance across four internal production datasets.
  • Automatically detects language and handles mid-recording switches in a single pass for dozens of languages.

The update will soon become the default in the Speech-to-Text API, with version 1.0 scheduled for deprecation in the coming weeks.