Alibaba drops five Qwen-Audio-3.1 voice models, cuts ASR pricing by 95%
aigclink · x · 2026-09-23
Alibaba released five Qwen-Audio-3.1 voice models with steep price cuts: ASR down 95%, Realtime down 85%, TTS down 70%, spanning understanding, generation, and real-time interaction.
- ASR: 30 languages + 16 Chinese dialects, native transcript cleanup, speaker-diarized output with timestamps, 160ms first-token streaming latency
- ASR-Next: goes beyond transcription to interpret emotion, ambient sounds, and answer audio questions
- TTS: cross-lingual voice consistency with emotion and speed control
- TTS-Next: generates complete audio works (voice + sound effects + ambience) at 48kHz, targeting podcasts, audiobooks, film, and games
- Realtime: full-duplex, interruptible, emotion-aware, seamless language switching, and can invoke agent tools mid-conversation
The author argues real-time voice is shifting from a "feature" to an "interaction entry point", with implications for customer service, assistants, and screenless hardware like AI glasses.
Related event: Alibaba Releases Five Qwen-Audio-3.1 Voice Models with Major Price Cuts(4 posts)→
More from Models
- JevBench hits HN frontpage, critics allege Jev-class model is a thin Qwen wrapper — airesearch12 · 2026-09-23
- Opus 5.5 One-Shots a Full Prince of Persia Level With Graphics, NPCs and Music — iannuttall · 2026-09-23
- Opus 5.5 Cut Prices 40% and Within a Day It Was Porting C to Rust: 8 Use Cases — alex_verem · 2026-09-23
- Beff Jezos jokes Opus 5.5 is 'post-slop', freeing readers from AI sludge prose — beffjezos · 2026-09-23
- Opus 5.5 One-Shots a Full Prince of Persia Level with NPCs, Sound and Music — iannuttall · 2026-09-23
- GPT-6 Sol Underperforms GPT-5.6 Max on DeepSWE, 68.8% vs 72.7% — banaxi-tech · 2026-09-23