Microsoft's MAI-Transcribe-2-Streaming claims the top speech-to-text accuracy spot

As seen on the 24/7 Wall St. homepage on October 1, 2026.

Microsoft taking the top speech-to-text benchmark spot with its own in-house model is one more piece of the stack it no longer needs OpenAI for.

Microsoft AI has released MAI-Transcribe-2-Streaming, taking the #1 spot for Final Transcript accuracy and First Partial Transcript accuracy on AA-WER Streaming with 2.5% WER at 0.13s after end of speech MAI-Transcribe-2-Streaming is @MicrosoftAI's new streaming Speech to Text https://t.co/Hv45RkabDB
  • Replies3
  • Reposts5
  • Likes55
Continue ReadingShow less

Microsoft AI released MAI-Transcribe-2-Streaming on October 1, and it immediately took the number one position for both Final Transcript accuracy and First Partial Transcript accuracy on the AA-WER Streaming benchmark, posting a 2.5% word error rate.

Those two metrics cover the full range of what a real-time transcription system needs to get right, and owning both top spots simultaneously is a meaningful result.

The model is built in-house at Microsoft AI. Every layer of the AI stack that Microsoft builds and benchmarks at the top level is one it no longer depends on an outside supplier to provide.

For a company that has spent billions establishing an AI partnership ecosystem, quietly releasing a streaming speech model that outperforms the field on a third-party benchmark signals that its internal capabilities are maturing faster than the public narrative has reflected.

Mentioned: MSFT