Alibaba launches Qwen-Audio-3.1: five models covering full audio stack from ASR to TTS

Scobleizer · x · 2026-09-28

Alibaba's Tongyi speech team released Qwen-Audio-3.1 with fully upgraded ASR, TTS and Realtime models, plus two new entries: TTS-Next for audio creation and ASR-Next for audio understanding. The five-model lineup forms a complete audio stack spanning understanding, generation, interaction and creation.

Original post →

More from Multimodal

Multimodal channel →