Alibaba Releases Qwen-Audio-3.0-ASR: 95% Accuracy in Medical Terms

智东西 · wechat · 2026-07-31

Alibaba's Qwen team released Qwen-Audio-3.0-ASR-Flash, a new speech recognition model. The update focuses on context consistency, industry-specific vocabulary, and custom hotwords. It achieves a 95.36% hit rate for medical terminology and over 90% accuracy in industrial and IT programming contexts.

Beyond raw accuracy, the model features built-in speech refinement capabilities. It automatically removes filler words, stutters, and self-corrections to output structured, written-quality text directly, matching the performance of traditional two-step pipelines. Additionally, it supports mixed recognition across 30 languages with an average semantic error rate of 17.09%. The model is now available via Alibaba Cloud's Bailian platform.

Related event: Alibaba Releases Qwen-Audio-3.0-ASR-Flash Model(3 posts)→

Original post →

More from Models

Models channel →