Qwen-Audio-3.0-TTS launches with 16 languages and better voice cloning
airesearch12 · x · 2026-07-21
Qwen releases Qwen-Audio-3.0-TTS in two variants:
- Flash for real-time interaction
- Plus for higher-quality generation
Notable additions include support for 16 languages, natural-language style control, fine-grained tags for non-verbal details, and more robust voice cloning from imperfect audio.
Related event: Alibaba Releases Qwen-Audio-3.0-TTS Voice Synthesis Model(5 posts)→
More from Multimodal
- CapCut demo mashes up Buddha, Sun Wukong, Thor and Ganesha in one video — AIandDesign · 2026-07-23
- AI Film 'Pomegranate' Set to Premiere Soon — Uncanny_Harry · 2026-07-23
- Canvas-to-Image turns identities, poses, and boxes into one RGB canvas — CSProfKGD · 2026-07-23
- Ambit v0.9.0 adds experimental Linux and macOS support for local AI images — Astra_Origin · 2026-07-23
- A Krea2 outpainting Space is trending on Hugging Face — yijunwang2 · 2026-07-23
- Paper: Video Generation Models are General-Purpose Vision Learners — dl_weekly · 2026-07-23