Real-Time STT WER Comparison: Three Major Models Evaluated

ArtificialAnlys · x · 2026-07-06

In partial transcription of the first sentence, Universal-3.5 Pro Realtime's Max Accuracy mode achieves 4.1% WER 0.44 seconds after speech ends. Its accuracy is second only to ElevenLabs Scribe v2 Realtime (3.6%, 0.13s) and slightly better than Cartesia Ink-2 outer endpoint (4.3%, 0.07s), though with slower text output. The Min Latency mode returns the first sentence in 0.39s with 6.1% WER, trading accuracy for lower latency.

Related event: AssemblyAI Launches Universal-3.5 Speech Models(4 posts)→

Original post →

More from Multimodal

Multimodal channel →