MAI-Transcribe-2-Streaming tops streaming STT leaderboard with 2.5% WER

ArtificialAnlys · x · 2026-10-02

On Final Transcript, Microsoft's MAI-Transcribe-2-Streaming hits 2.5% WER at 0.13s after end of speech, ranking #1 of 38 models on Artificial Analysis's AA-WER Streaming leaderboard. It beats Grok Voice Transcribe 2.0 (2.7% WER, 0.49s), Muse Voice Transcribe (3.1%, 0.16s), Cartesia Ink Preview with external endpoints (3.1%, 0.11s) and ElevenLabs Scribe v2 Realtime (3.6%, 0.14s), and is faster than all except Cartesia — sitting on the low-error end of the accuracy-vs-latency Pareto frontier.

Related event: Microsoft's MAI-Transcribe-2-Streaming Tops Speech Transcription Benchmarks(2 posts)→

Original post →

More from Models

Models channel →