Alibaba Releases Qwen-Audio-3.0-TTS Voice Synthesis Model
Alibaba's Tongyi Lab officially released the Qwen-Audio-3.0-TTS voice synthesis model. It focuses on multilingual support and highly controllable voice generation, lowering the latency barrier for real-time voice interaction.
Versions and Core Parameters
The model comes in two versions: the Flash version for real-time interaction with a first-packet latency of around 300ms, and the Plus version for higher-quality generation with better naturalness and timbre restoration.
Multilingual and Voice Capabilities
Qwen-Audio-3.0-TTS supports 16 languages and 20 dialects. By providing just a single reference audio clip, the model can speak across different languages while maintaining the exact same voice timbre.
Controllability and Style Instructions
The model features strong controllability. Users can control the tone and style using natural language, and embed fine-grained tags like [gasp] directly into the text for precise emotional and acoustic details.
2026-07-20 ~ 2026-07-21 · 5 related posts
- Episode 1: Alibaba's Qwen-Audio-3.0-TTS-Plus Tops Speech Arena(2026-07-15, 2 posts)
- Episode 2: Alibaba Releases Qwen-Audio-3.0-TTS Voice Synthesis Model(2026-07-20, 5 posts)
- Episode 3: Alibaba's Qwen TTS Tops Charts, API Access Only(2026-07-21, 3 posts)
Primary sources
- [source] Alibaba Qwen launches Qwen-Audio-3.0-TTS for real-time and premium speech — 千问大模型 · 2026-07-20
- Qwen-Audio-3.0-TTS launches with 16 languages and better voice cloning — airesearch12 · 2026-07-21
- [source] Qwen-Audio-3.0-TTS speaks 16 languages and 20 dialects from one reference voice — xiaohu · 2026-07-21