Qwen-Audio-3.0-TTS-Plus Tops Speech Leaderboard
ArtificialAnlys · x · 2026-07-15
Alibaba's Qwen-Audio-3.0-TTS-Plus took first place on Artificial Analysis's Speech Arena 'Provider Voices' leaderboard with an Elo of 1236, slightly edging out Simba 3.2 (1234) and beating Gemini 3.1 Flash TTS and Sonic 3.5.
While this model delivers more natural pronunciation and context-aware intonation, it falls short on speed: generating 16 characters per second, notably slower than Simba 3.2 (30.2), Gemini 3.1 Flash TTS (27), and Sonic 3.5 (120).
Priced at $27.59 / 1 million characters via Alibaba Cloud Model Studio, it sits between Gemini 3.1 Flash TTS ($18.31) and Sonic 3.5 ($39.00), and is more expensive than Simba 3.2 ($10.00).
Related event: Alibaba's Qwen-Audio-3.0-TTS-Plus Tops Speech Arena(2 posts)→
More from Multimodal
- Gemini Omni Flash turns a boat cabin into a cave in Flow by Google — chrisfirst · 2026-07-22
- A simple workflow to turn a photo into an image prompt using Gemini, Grok, or GPT Image — harshitagu72595 · 2026-07-22
- A Reddit user proposes a consistency LoRA to keep anime and game scenes visually stable — ThirdWorldBoy21 · 2026-07-22
- Hand-painted figurines run through Seedance look eerily alive — cocktailpeanut · 2026-07-22
- An AI agent-made bayou country music video is making the rounds on Reddit — LazyKaleidoscope4696 · 2026-07-22
- Testing Qwen 3 Image: Map Borders Shift Based on Prompts, Includes Chinese Labels — NirantK · 2026-07-22