FULL STORY

Alibaba's Qwen-Audio-3.0-TTS: From Topping Charts to Official Release

Alibaba's Qwen-Audio-3.0-TTS first topped the Speech Arena leaderboard and was subsequently officially released, featuring multilingual support and low-latency real-time interaction.

2026-07-15 ~ 2026-07-21 · 2 episodes · 7 posts

Episode 1 · Alibaba's Qwen-Audio-3.0-TTS-Plus Tops Speech Arena (2026-07-15, 2 posts)

Artificial Analysis updated its Speech Arena, releasing voting and sample exploration. Alibaba's Qwen-Audio-3.0-TTS-Plus topped the Provider Voices leaderboard for its quality, despite lagging in generation speed.

Episode 2 · Alibaba Releases Qwen-Audio-3.0-TTS Speech Synthesis Model (2026-07-20, 5 posts)

Alibaba Tongyi Lab released the Qwen-Audio-3.0-TTS speech synthesis model, featuring multilingual support and highly controllable voice generation, further reducing latency for real-time voice interaction.

Versions and Key Parameters

The model comes in two versions: Flash for real-time interaction with ~300ms first-packet latency, and Plus for higher-quality generation with better naturalness and timbre fidelity.

Multilingual and Voice Capabilities

Qwen-Audio-3.0-TTS supports 16 languages and 20 dialects. Users only need a reference audio clip, and the model can speak cross-lingually while maintaining the same voice.

Controllability and Style Instructions

The model offers strong controllability. Users can control tone and style via natural language, and embed fine-grained tags like `[gasp]` within text for precise emotional and vocal detail.