SpaceXAI has released Grok Voice Transcribe 2.0, a speech-to-text model available via hosted API that claims twice the accuracy of version 1.0 at the same price point. The model targets difficult audio conditions such as noisy phone lines and competing voices, running in both batch and real-time streaming modes.
- Ranked first among 32 streaming models on the Artificial Analysis leaderboard using the AA-WER Streaming benchmark.
- Achieved a word error rate drop from 20.6% to 6.8% on short phrases across 19 languages, representing roughly 67% fewer errors.
- Supports automatic language detection and mid-recording language switching for dozens of languages.
- Includes features like speaker diarization, word-level timestamps, key term biasing, and filler word removal at no extra cost.
- Priced at $0.10 per hour for batch transcription and $0.20 per hour for streaming.
Atlassian Loom has adopted the API to transcribe every video in its workflow, citing higher accuracy than its previous solution.