Alibaba launches Qwen-Audio-3.0-TTS with 16 languages and 3-minute one-pass audio
Alibaba_Qwen · x · 2026-07-23
Alibaba’s Qwen team introduced Qwen-Audio-3.0-TTS, a new text-to-speech model with two variants: Flash for real-time interaction and Plus for higher-quality generation.
Key upgrades include:
- Fine-grained inline control tags such as [whisper], [angry], [breaths], and [laughs]
- Natural-language steering like “read this slowly, like a bedtime story”
- Support for 16 languages
- Better output from noisy reference audio
- One-pass long-form generation up to 3 minutes
Alibaba says the model is now #1 on the Artificial Analysis TTS Leaderboard. The post links to the blog and API.
More from Multimodal
- Open-source node-based LoRA trainer puts captioning, checkpoints and VRAM stats in one graph — ashishsanu · 2026-07-23
- TERRA-129 debuts as an AI-animated sci-fi episode credited to Matygoo — Matygoo1 · 2026-07-23
- Kling AI is said to handle close-up facial expressions better — burny_tech · 2026-07-23
- FameGrid Krea 2 aims to generate more realistic social-media-style images — UltraMuseArt · 2026-07-23
- A reusable ChatGPT image prompt for a realistic portrait plus doodle-shadow twin — SimplyAnnisa · 2026-07-23
- User showcases LTX 2.3 animations with a cinematic ogre-at-the-cake scene — Wise_Revolution385 · 2026-07-23