TTS benchmark breakdown: exact-sequence category shows 5.3%-85.7% model spread
ArtificialAnlys · x · 2026-09-22
Artificial Analysis published category-level results for its pronunciation robustness benchmark: Gemini 3.1 Flash TTS leads contextually appropriate (96.4%) and expanding shorthand (84.4%, barely ahead of SpaceXAI TTS at 84.3%); SpaceXAI TTS tops exact-sequence preservation at 85.7%, 5.8 points over Realtime TTS-2 (79.9%); Qwen-Audio-3.0-TTS-Plus leads standalone terms (95.5%). The widest model spread is in exact sequences, from 5.3% to 85.7%.
More from Multimodal
- Striking human-motion visualizations made with Sentinel — Kyrannio · 2026-09-22
- New Pika impresses early users as the video startup returns to form — Kyrannio · 2026-09-22
- Tencent's Hunyuan Image 3.5 lands on OnSolo: 5 refs, 2K output, 1 credit — Div_pradeep · 2026-09-22
- Museum statue meme generated with AI — prompt included — umesh_ai · 2026-09-22
- Princeton's VideoGen-Agent Adds 19 Points via Agentic RL Tool Use for Video — princetonu · 2026-09-22
- Intel Ships Day-0 OpenVINO Support for Qwen-Image-2.1 — Alibaba_Qwen · 2026-09-22