Alibaba launches Qwen-Audio-3.1: five models covering full audio stack from ASR to TTS
Scobleizer · x · 2026-09-28
Alibaba's Tongyi speech team released Qwen-Audio-3.1 with fully upgraded ASR, TTS and Realtime models, plus two new entries: TTS-Next for audio creation and ASR-Next for audio understanding. The five-model lineup forms a complete audio stack spanning understanding, generation, interaction and creation.
More from Multimodal
- Same Prompt, 4x Difference: Claude Builds Three.js Output in 2 Hours vs 8 for Unreal — nptacek · 2026-09-28
- Udio Has Far More Range but Fails Harder, While Suno Stays in Narrow Lanes — gandamu_ml · 2026-09-28
- AI product shots finally nail cream texture that used to look like slime — Stock_Appeal_8654 · 2026-09-28
- Open-source ClaudeAnimationBase kit animates hand-painted cartoons with Claude, ships 31 emotions — FinanceYF5 · 2026-09-28
- Seedance 2.5 nails early-2000s DV camcorder look in ultra-realistic AI video — SimplyAnnisa · 2026-09-28
- Two ComfyUI Nodes Fix Pixel Drift in Qwen 2.1 Image Editing — Technical_Fish_9638 · 2026-09-28