FULL STORY

Alibaba's Qwen TTS Tops Speech Arena

Alibaba released the Qwen-Audio-3.0-TTS model. Its Plus version topped the Speech Arena, demonstrating advanced multilingual and low-latency capabilities.

2026-07-15 ~ 2026-07-22 · 3 episodes · 10 posts

Episode 1 · Alibaba's Qwen-Audio-3.0-TTS-Plus Tops Speech Arena (2026-07-15, 2 posts)

Artificial Analysis updated its Speech Arena, releasing voting and sample exploration. Alibaba's Qwen-Audio-3.0-TTS-Plus topped the Provider Voices leaderboard for its quality, despite lagging in generation speed.

Episode 2 · Alibaba Releases Qwen-Audio-3.0-TTS Voice Synthesis Model (2026-07-20, 5 posts)

Alibaba's Tongyi Lab officially released the Qwen-Audio-3.0-TTS voice synthesis model. It focuses on multilingual support and highly controllable voice generation, lowering the latency barrier for real-time voice interaction.

Versions and Core Parameters

The model comes in two versions: the Flash version for real-time interaction with a first-packet latency of around 300ms, and the Plus version for higher-quality generation with better naturalness and timbre restoration.

Multilingual and Voice Capabilities

Qwen-Audio-3.0-TTS supports 16 languages and 20 dialects. By providing just a single reference audio clip, the model can speak across different languages while maintaining the exact same voice timbre.

Controllability and Style Instructions

The model features strong controllability. Users can control the tone and style using natural language, and embed fine-grained tags like [gasp] directly into the text for precise emotional and acoustic details.

Episode 3 · Alibaba's Qwen TTS Tops Charts, API Access Only (2026-07-21, 3 posts)

Alibaba's Qwen-Audio-3.0-TTS-Plus recently topped the Artificial Analysis Speech Arena leaderboard, beating Gemini. However, the model is currently restricted to API access without an open-weight release, sparking discussions about its commercialization strategy.