Real-Time STT WER Comparison: Three Major Models Evaluated
ArtificialAnlys · x · 2026-07-06
In partial transcription of the first sentence, Universal-3.5 Pro Realtime's Max Accuracy mode achieves 4.1% WER 0.44 seconds after speech ends. Its accuracy is second only to ElevenLabs Scribe v2 Realtime (3.6%, 0.13s) and slightly better than Cartesia Ink-2 outer endpoint (4.3%, 0.07s), though with slower text output. The Min Latency mode returns the first sentence in 0.39s with 6.1% WER, trading accuracy for lower latency.
Related event: AssemblyAI Launches Universal-3.5 Speech Models(4 posts)→
More from Multimodal
- Midjourney style code share: --sref 2912175708 — tisch_eins · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- MiniMax H3 MAX nails cooking anime clips: 15-second curry demo with prompts shared — Hailuo_AI · 2026-09-11
- MiniMax Music Production Toolkit 2.5 for ComfyUI adds full mastering chain — Vivid_Promise1700 · 2026-09-11
- New Node Finder for ComfyUI ranks fresh nodes by star velocity and recency — Luke2642 · 2026-09-11
- Using a finisher move on one mosquito with MiniMax H3 MAX — the bug survives — Hailuo_AI · 2026-09-11