Qwen-Audio-3.0-TTS-Plus Tops Speech Leaderboard
ArtificialAnlys · x · 2026-07-15
Alibaba's Qwen-Audio-3.0-TTS-Plus took first place on Artificial Analysis's Speech Arena 'Provider Voices' leaderboard with an Elo of 1236, slightly edging out Simba 3.2 (1234) and beating Gemini 3.1 Flash TTS and Sonic 3.5.
While this model delivers more natural pronunciation and context-aware intonation, it falls short on speed: generating 16 characters per second, notably slower than Simba 3.2 (30.2), Gemini 3.1 Flash TTS (27), and Sonic 3.5 (120).
Priced at $27.59 / 1 million characters via Alibaba Cloud Model Studio, it sits between Gemini 3.1 Flash TTS ($18.31) and Sonic 3.5 ($39.00), and is more expensive than Simba 3.2 ($10.00).
Related event: Alibaba's Qwen-Audio-3.0-TTS-Plus Tops Speech Arena(2 posts)→
More from Multimodal
- Tencent Hunyuan releases AuK code and weights on GitHub with ComfyUI and fine-tuning support — aigclink · 2026-09-11
- Tencent open-sources AuK, a unified 1.5B speech generation and editing model — aigclink · 2026-09-11
- Creator turns Bahamut vs Tiamat rivalry into an AI cinematic battle with Midjourney, GPT Image 2 and Seedance — azed_ai · 2026-09-11
- invideo launches AI agent-powered editor to automate repetitive editing tasks — azed_ai · 2026-09-11
- fable 5.1 recreates The Starry Night with 256,157 JavaScript brush strokes — cedric_chee · 2026-09-11
- GPT-6 Astra + Hyper3D Rodin MCP Generates 3D Assets in One Agent Flow — ahuja_priyank · 2026-09-11