Alibaba launches Qwen-Audio-3.0-TTS with 16 languages and 3-minute one-pass audio
Alibaba_Qwen · x · 2026-07-23
Alibaba’s Qwen team introduced Qwen-Audio-3.0-TTS, a new text-to-speech model with two variants: Flash for real-time interaction and Plus for higher-quality generation.
Key upgrades include:
- Fine-grained inline control tags such as [whisper], [angry], [breaths], and [laughs]
- Natural-language steering like “read this slowly, like a bedtime story”
- Support for 16 languages
- Better output from noisy reference audio
- One-pass long-form generation up to 3 minutes
Alibaba says the model is now #1 on the Artificial Analysis TTS Leaderboard. The post links to the blog and API.
More from Multimodal
- Lumara AI Film Festival Comes to NYC Oct 26, Top AI Filmmakers to Compete — 0xAllen_ · 2026-09-11
- Pterodactyl Detective: An AI-Generated Proof-of-Concept Trailer — PterodactylDetective · 2026-09-11
- Imperium Game Trailer Showcases AI Video Generation — keaslenyt · 2026-09-11
- FLUX.2 Klein Drifts Hard on Character Expressions While Free Gemini Holds Likeness — wacomlover · 2026-09-11
- Tencent Hunyuan releases AuK code and weights on GitHub with ComfyUI and fine-tuning support — aigclink · 2026-09-11
- Creator turns Bahamut vs Tiamat rivalry into an AI cinematic battle with Midjourney, GPT Image 2 and Seedance — azed_ai · 2026-09-11