Alibaba Releases Qwen-Audio-3.0-TTS Voice Synthesis Model

Alibaba's Tongyi Lab officially released the Qwen-Audio-3.0-TTS voice synthesis model. It focuses on multilingual support and highly controllable voice generation, lowering the latency barrier for real-time voice interaction.

Versions and Core Parameters

The model comes in two versions: the Flash version for real-time interaction with a first-packet latency of around 300ms, and the Plus version for higher-quality generation with better naturalness and timbre restoration.

Multilingual and Voice Capabilities

Qwen-Audio-3.0-TTS supports 16 languages and 20 dialects. By providing just a single reference audio clip, the model can speak across different languages while maintaining the exact same voice timbre.

Controllability and Style Instructions

The model features strong controllability. Users can control the tone and style using natural language, and embed fine-grained tags like [gasp] directly into the text for precise emotional and acoustic details.

2026-07-20 ~ 2026-07-21 · 5 related posts

Full story(3 episodes)→

Primary sources

2 near-duplicate retellings: 通义实验室 · 智东西