Alibaba's Qwen Launches Five-Model Audio Stack, Slashes TTS 70% and ASR up to 95%

Alibaba_Qwen · x · 2026-09-23

Alibaba's Qwen team announced Qwen-Audio-3.1: fully upgraded ASR, TTS, and Realtime models, plus two newcomers — TTS-Next for audio creation and ASR-Next for audio understanding — forming a five-model stack covering understanding, generation, interaction, and creation. Pricing drops sharply: TTS 70% off, Realtime 85% off, ASR up to 95% off.

Highlights:

Related event: Qwen Launches Five Audio Models with up to 95% Price Cut(4 posts)→

Original post →

More from Multimodal

Multimodal channel →