Qwen Releases Qwen-Audio-3.0-ASR: Supports Long-Context and 30 Languages

千问大模型 · wechat · 2026-07-31

Alibaba's Qwen has officially released the Qwen-Audio-3.0-ASR-Flash speech recognition model, focusing on solving professional vocabulary recognition and long-audio consistency. It previously ranked first globally on ArtificialAnalysis with a 1.7% Character Error Rate.

The new model introduces four major upgrades:

Additionally, the Streaming version for low-latency scenarios reduces the Chinese error rate in complex industrial contexts to 7.8% while maintaining a 300ms latency. The models are now available on Alibaba Cloud's Bailian platform.

Original post →

More from Models

Models channel →