Tencent Hunyuan Launches HyASR3.0: Reduces Multilingual WER to Around 3%

腾讯混元 · wechat · 2026-08-04

Tencent Hunyuan officially released HyASR3.0preview, its new generation speech recognition model. Based on the Hy3 LLM, the model adopts a MoE architecture and a self-developed unsupervised speech encoder, integrating high-precision speech recognition with deep semantic understanding.

According to official data, HyASR3.0preview keeps the Word Error Rate (WER) for multiple languages at around 3% across various open-source evaluation sets (Mandarin 3.34%, English 2.62%, Cantonese 3.12%). The model focuses on improving general recognition accuracy, contextual intent understanding, professional scene adaptation (supporting hotword injection), and stability in complex acoustic environments (high noise, whispers). The model is now available via API on Tencent Cloud and is integrated into the Yuanbao app.

Related event: Tencent Hunyuan Releases Next-Gen HyASR3.0 Speech Recognition Model(2 posts)→

Original post →

More from Models

Models channel →