Qwen-Audio-3.1 models go live: ASR-Flash priced at ¥0.8/M input tokens
aigclink · x · 2026-09-23
Follow-up on the Qwen-Audio-3.1 launch: ASR, TTS, TTS-Next, and Realtime APIs are already live, with ASR-Next coming soon.
The model page for Qwen-Audio-3.1-ASR-Flash shows short-audio recognition optimized for multilingual content, Chinese dialects, and classical Chinese prosody, with industrial-grade speaker diarization. Pricing is ¥0.8/M input tokens and ¥2.7/M output tokens, with a 600 RPM rate limit, plus function calling, structured output, and batch tasks.
Related event: Qwen Launches Five Audio Models with up to 95% Price Cut(4 posts)→
More from Multimodal
- Dev builds a character-swap LoRA dataset end-to-end with Codex and GPT Image 2.5 — ostrisai · 2026-09-23
- Tencent Hunyuan Image3.5 preview went live Sept 22, free for a limited time — HeyAmit_ · 2026-09-23
- Hands-on with Tencent's Hy Image 3.5: strong text rendering, editing and 5-image references — HeyAmit_ · 2026-09-23
- ImIR replaces text prompts with image instructions for all-in-one restoration — Süleyman Aslan · 2026-09-23
- Alibaba's Qwen Launches Five-Model Audio Stack, Slashes TTS 70% and ASR up to 95% — Alibaba_Qwen · 2026-09-23
- Microsoft ships MAI-Transcribe 2 with multilingual speech, diarization, timestamps — lee_stott · 2026-09-23