Microsoft AI has released MAI-Transcribe-2-Streaming, its first streaming speech-to-text model, which ranks #1 of 38 models on the Artificial Analysis AA-WER Streaming index. Launched on October 1, 2026, the model transcribes 60 languages with continuous automatic language detection and targets voice agents and live captions.

  • Achieves 2.5% Word Error Rate (WER) for both final transcripts (at 0.13s latency) and first partial transcripts (at 0.12s latency).
  • Emits initial hypotheses just over 100ms after receiving audio, allowing agents to reason mid-sentence.
  • Costs $0.54 per hour during an introductory period through the end of 2026.
  • Available via Realtime API (WebSocket) and Azure Speech SDK, with integration in MAI Playground and Vercel.

The model's ability to provide partial transcripts as accurate as final ones enables real-time interaction for applications where latency is critical.