TTS benchmark breakdown: exact-sequence category shows 5.3%-85.7% model spread

ArtificialAnlys · x · 2026-09-22

Artificial Analysis published category-level results for its pronunciation robustness benchmark: Gemini 3.1 Flash TTS leads contextually appropriate (96.4%) and expanding shorthand (84.4%, barely ahead of SpaceXAI TTS at 84.3%); SpaceXAI TTS tops exact-sequence preservation at 85.7%, 5.8 points over Realtime TTS-2 (79.9%); Qwen-Audio-3.0-TTS-Plus leads standalone terms (95.5%). The widest model spread is in exact sequences, from 5.3% to 85.7%.

Related event: Artificial Analysis Launches Pronunciation Robustness Benchmark for TTS Models(6 posts)→

Original post →

More from Multimodal

Multimodal channel →