Qwen-Audio-3.1 models go live: ASR-Flash priced at ¥0.8/M input tokens

aigclink · x · 2026-09-23

Follow-up on the Qwen-Audio-3.1 launch: ASR, TTS, TTS-Next, and Realtime APIs are already live, with ASR-Next coming soon.

The model page for Qwen-Audio-3.1-ASR-Flash shows short-audio recognition optimized for multilingual content, Chinese dialects, and classical Chinese prosody, with industrial-grade speaker diarization. Pricing is ¥0.8/M input tokens and ¥2.7/M output tokens, with a 600 RPM rate limit, plus function calling, structured output, and batch tasks.

Related event: Qwen Launches Five Audio Models with up to 95% Price Cut(4 posts)→

Original post →

More from Multimodal

Multimodal channel →