HeyGen Voice scores 83.1% on pronunciation robustness, ranking 10th of 29 TTS models
ArtificialAnlys · x · 2026-10-10
Artificial Analysis' Pronunciation Robustness benchmark, where human reviewers judge whether TTS models correctly pronounce challenging text, puts HeyGen Voice at 83.1% overall — 10th of 29 models.
- Preserving exact sequences: 83.3%, #2 behind SpaceXAI TTS (85.7%)
- Standalone terms: 90.5%, led by Eleven v3 Conversational (95.1%)
- Contextually appropriate: 89.8%, led by Gemini 3.8 Flash TTS (97.9%)
- Expanding shorthand: 77.0%, below Qwen-Audio-3.1-TTS-Plus (80.2%); Eleven v4 leads at 94.1%
More from Models
- Anthropic's refusal of sadistic-persona requests sparks community backlash debate — mcraddock · 2026-10-10
- Integer multiplication tracker for LLMs adds zoom to visualize recent progress — RexDouglass · 2026-10-10
- lateinteraction keynote take: stronger models will need longer prompts, not shorter — lateinteraction · 2026-10-10
- Cloudflare ships open-weight Clef-omni with audio/video input, cuts Clef-flash to $0.038/M tokens — Cloudflare Blog · 2026-10-10
- Radiologist pressures AI three times to sign off a benign breast report — it holds firm — FellMentKE · 2026-10-10
- Radiologist tests medical AI: model refuses wrong BI-RADS and demands more evidence — FellMentKE · 2026-10-10