Alibaba launches Qwen-Audio-3.0-TTS with 16 languages and 3-minute one-pass audio

Alibaba_Qwen · x · 2026-07-23

Alibaba’s Qwen team introduced Qwen-Audio-3.0-TTS, a new text-to-speech model with two variants: Flash for real-time interaction and Plus for higher-quality generation.

Key upgrades include:

Alibaba says the model is now #1 on the Artificial Analysis TTS Leaderboard. The post links to the blog and API.

Original post →

More from Multimodal

Multimodal channel →