Alibaba's Qwen launches 5-model audio stack Qwen-Audio-3.1, cuts TTS prices ~70%
airesearch12 · x · 2026-09-23
Alibaba's Qwen team launched Qwen-Audio-3.1, a five-model audio stack covering understanding, generation, interaction and creation:
- ASR/ASR-Next: stronger multilingual and dialect recognition plus native polishing that strips fillers and repetitions; ASR-Next adds multi-speaker labels, timestamps, and sound/event understanding
- TTS/TTS-Next: aimed at audio creation
- Realtime: fully upgraded real-time speech interaction
Pricing was cut across the lineup: 70% off TTS, 85% off Realtime, and up to 95% off ASR. Open-weight release remains unconfirmed.
More from Models
- JevBench v1.4 follow-up link: details of the anti-benchmaxxing methodology — airesearch12 · 2026-09-23
- Casually mentioning a cat makes Opus 5.5 grill the user for cat facts — voooooogel · 2026-09-23
- JevBench v1.4 adds 308 evolving sealed tasks to stop benchmaxxing, now covers 70+ models — airesearch12 · 2026-09-23
- Stop benchmarking LLMs with 3D games, says Abacus.AI CEO — labs fine-tune for it — bindureddy · 2026-09-23
- Opus 5.5 impresses with a hand-crafted Mona Lisa SVG — djcows · 2026-09-23
- MazeBench 3D Spatial Benchmark: Opus 5.5 Hits 6% While GPT-6 Sol and Grok 4.7 Score Just 1% — daniel_mac8 · 2026-09-23