New TTS pronunciation benchmark: Gemini 3.1 Flash TTS leads at 88.1%
ArtificialAnlys · x · 2026-09-22
Artificial Analysis launched a Pronunciation Robustness benchmark measuring how reliably TTS models say challenging text, using 454 sentences with 701 target words across four categories (context-dependent readings, shorthand expansion, exact sequences, standalone terms).
- Overall: Google Gemini 3.1 Flash TTS leads at 88.1%, followed by SpaceXAI TTS (87.6%), ElevenLabs Eleven v3 (85.6%), Eleven v3 Conversational (84.8%), and Alibaba Qwen-Audio-3.0-TTS-Plus (81.6%).
- By category: Gemini leads contextually appropriate (96.4%) and expanding shorthand (84.4%); SpaceXAI TTS tops exact-sequence preservation (85.7%); Qwen leads standalone terms (95.5%). Exact sequences show the widest model spread (5.3%–85.7%).
- Hardest categories: expanding shorthand (62.4%) and exact sequences (62.9%) trail the others; input normalization should lift scores considerably.
- Preference ≠ accuracy: Sonic 3.6 ranks #1 on the Provider Voice Arena (1276 Elo) but #11 on pronunciation (74.5%), while Gemini 3.1 Flash ranks #9 on the Arena yet #1 here.
The metric matters for production voice agents that must read names, account details and amounts correctly.
More from Multimodal
- BUPT study: RoPE attention decay causes video diffusion models to violate physics — BUPT-CIST · 2026-09-22
- Tencent ARC's WorldCrafter adds implicit 3D-aware memory to video world models — TencentARC · 2026-09-22
- Grok 4.7 made this in Blender — demo shows the model driving 3D software — iamfakhrealam · 2026-09-22
- Kyutai releases Voice of Reason, a speech-native reasoning model hitting 77.1% on GSM8K — alexcovo_eth · 2026-09-22
- Qwen-Image local on a 24GB MacBook Pro takes 5-6 minutes per image — vista8 · 2026-09-22
- Tencent Hunyuan ships Hy Image3.5 preview: +30% human eval win rate at $0.024 per image — TencentHunyuan · 2026-09-22