Eleven v4 Tops TTS Arena, Tops Pronunciation Benchmark at $80/1M Chars
ArtificialAnlys · x · 2026-09-28
Artificial Analysis evaluated ElevenLabs' new TTS model Eleven v4:
- #1 on Provider Voice with Elo 1,319 across 1,674 appearances, ahead of Cartesia Sonic 3.6 (1,276) and Google Gemini 3.8 Flash TTS (1,267); first place in all four categories.
- #2 on Controlled Voice (Elo 1,157), behind only Alibaba's Qwen-Audio-3.1-TTS-Plus (1,178) and well above Eleven v3 (1,073).
- Pronunciation Robustness 91.7%, the highest ever measured, vs 89.5% for Gemini 3.8 Flash TTS and 85.6% for v3.
- Language support expands from 70+ to 90+ languages.
- Priced at $80/1M chars vs Sonic 3.6's $49 and Gemini 3.8 Flash TTS's $16.49; speed is 73.4 chars/sec (v3: 42.5).
Related event: ElevenLabs Launches Eleven v4 and v4 Turbo, Topping TTS Leaderboards(11 posts)→
More from Multimodal
- User has Opus 5.5 make a space exploration film — every frame and note written in code — prasenx · 2026-09-28
- Dev laments Opus 5.5 can one-shot music videos after building his own pipeline — nptacek · 2026-09-28
- Kling 4.0 Confirmed for October; Flash Model Live Now for Ultra Yearly Subscribers — koltregaskes · 2026-09-28
- Early hands-on with Kling 4.0 Flash: impressive prompt understanding — umesh_ai · 2026-09-28
- KP rehashes WebMCP basics with an explainer video generated by Opus 5.5 — thisiskp_ · 2026-09-28
- AI motion design's weak spot: music syncing and default sound effects — ciguleva · 2026-09-28