Alibaba's Qwen-Audio-3.1 ships five voice models, cuts ASR price 95%
千问大模型 · wechat · 2026-09-23
Alibaba's Qwen team released the Qwen-Audio-3.1 family: five models spanning understanding, generation, interaction, and creation — upgraded ASR, TTS, and Realtime, plus new ASR-Next (audio understanding) and TTS-Next (full audio creation) models. Prices dropped across the board: TTS 70%, Realtime 85%, ASR 95%.
Highlights
- ASR: 30 languages, 16 Chinese dialects (avg CER 10.38% on internal dialect set), native transcript polishing, speaker-labeled structured output, 160ms streaming first-token latency
- ASR-Next: general audio understanding — emotion, environmental and mechanical sounds, audio QA and event localization
- TTS: natural cross-language voice transfer, instruction-controlled emotion and pacing
- TTS-Next: unified generation of speech, sound effects, and ambience with timestamp control and 48kHz output
- Realtime: full-duplex conversation with interruptions, emotion/intent awareness, mid-dialogue tool calling
APIs are live on the Qwen platform and being integrated into agents like Qoder and AI glasses hardware.
Related event: Qwen Launches Five Audio Models with up to 95% Price Cut(4 posts)→
More from Multimodal
- Opus 5.5 one-shots a full newsletter launch video, creator says "it's over for video guys" — alex_verem · 2026-09-23
- PixVerse R2 hands-on: AI video that becomes a world you walk through with WASD — HeyAmit_ · 2026-09-23
- One-Person AI Music Video Took a Month: 20 Stills and 10 Takes Per Clip — kraussian · 2026-09-23
- MiniMax H3 Ref2VA Freezes at Model Initializing With an 8th Reference Image — itchplease · 2026-09-23
- H3 long-video degradation workaround: a second noise-injection refine stage in latent space — xyzdist · 2026-09-23
- Dev builds a character-swap LoRA dataset end-to-end with Codex and GPT Image 2.5 — ostrisai · 2026-09-23