Eleven v4 Tops New TTS Pronunciation Robustness Benchmark, Scoring Above 93% in Three Categories
ArtificialAnlys · x · 2026-09-28
Artificial Analysis introduced Pronunciation Robustness, a benchmark measuring whether TTS models correctly pronounce challenging text across four categories, with human reviewers judging each clip against pre-agreed accepted pronunciations.
- Expanding Shorthand: Eleven v4 ranks #1 at 94.1%, ahead of Gemini 3.8 Flash TTS at 86.1% and up from Eleven v3's 82.7%
- Contextually Appropriate: Eleven v4 scores 94.1%, behind Gemini 3.8 Flash TTS at 97.9%
- Standalone Terms: Eleven v4 scores 93.2%, with Eleven v3 Conversational leading at 95.1%
- Preserving Exact Sequences: Eleven v4 scores 78.8% (v3: 71.2%), with SpaceXAI TTS leading at 85.7%
Eleven v4 is the only model to score above 93% in three of the four categories.
More from Multimodal
- Horse gaits one-shot with Claude Opus 5.5: impressive coat shine and musculature — shekitup · 2026-09-28
- Ad Buyer Ditches Image Models, Uses Seedance 2.5 for Realistic AI UGC Characters — churchkey · 2026-09-28
- One-word Midjourney prompt: 'Pulchritudinous' at --ar 5:4 --v 8.2 — tisch_eins · 2026-09-28
- Prompt Template Makes Qwen Image 2.1 Design Like Closed-Source Models — wjc_5 · 2026-09-28
- A ChatGPT prompt turns GPT Image into an editorial fashion photographer — umesh_ai · 2026-09-28
- Synthesia ships Express-3, its best avatar model, free for users on all plans — synthesiaIO · 2026-09-28