Qwen-Audio-3.0-TTS adds finer voice control and broader language coverage

智东西 · wechat · 2026-07-20

Alibaba’s Qwen team has released Qwen-Audio-3.0-TTS, a new text-to-speech model with two variants:

The model adds finer control over tone, emotion, breathing, and style through structured tags such as [gasp], [giggles], and [angry], and it also supports free-form natural language instructions for character, mood, scene, and speaking rate.

Other reported upgrades:

Both Plus and Flash are now available on Alibaba Cloud Bailian.

Related event: Alibaba Releases Qwen-Audio-3.0-TTS Speech Synthesis Model(5 posts)→

Original post →

More from Multimodal

Multimodal channel →