Alibaba's Qwen Launches Five-Model Audio Stack, Slashes TTS 70% and ASR up to 95%
Alibaba_Qwen · x · 2026-09-23
Alibaba's Qwen team announced Qwen-Audio-3.1: fully upgraded ASR, TTS, and Realtime models, plus two newcomers — TTS-Next for audio creation and ASR-Next for audio understanding — forming a five-model stack covering understanding, generation, interaction, and creation. Pricing drops sharply: TTS 70% off, Realtime 85% off, ASR up to 95% off.
Highlights:
- ASR: stronger multilingual/dialect recognition, native polishing that strips fillers and repetitions for cleaner transcripts.
- ASR-Next: multi-speaker ASR with speaker labels and timestamps, plus emotion, ambient and machine-sound understanding for sound captioning, event localization, audio QA and reasoning.
- TTS: multilingual/dialect synthesis with natural cross-lingual voice transfer; emotion, speed, and style controllable via simple instructions.
- TTS-Next: unified LM + diffusion framework.
Related event: Qwen Launches Five Audio Models with up to 95% Price Cut(4 posts)→
More from Multimodal
- Opus 5.5 one-shots a full newsletter launch video, creator says "it's over for video guys" — alex_verem · 2026-09-23
- PixVerse R2 hands-on: AI video that becomes a world you walk through with WASD — HeyAmit_ · 2026-09-23
- One-Person AI Music Video Took a Month: 20 Stills and 10 Takes Per Clip — kraussian · 2026-09-23
- MiniMax H3 Ref2VA Freezes at Model Initializing With an 8th Reference Image — itchplease · 2026-09-23
- H3 long-video degradation workaround: a second noise-injection refine stage in latent space — xyzdist · 2026-09-23
- Dev builds a character-swap LoRA dataset end-to-end with Codex and GPT Image 2.5 — ostrisai · 2026-09-23