Tencent Hunyuan Launches HyASR3.0: Reduces Multilingual WER to Around 3%
腾讯混元 · wechat · 2026-08-04
Tencent Hunyuan officially released HyASR3.0preview, its new generation speech recognition model. Based on the Hy3 LLM, the model adopts a MoE architecture and a self-developed unsupervised speech encoder, integrating high-precision speech recognition with deep semantic understanding.
According to official data, HyASR3.0preview keeps the Word Error Rate (WER) for multiple languages at around 3% across various open-source evaluation sets (Mandarin 3.34%, English 2.62%, Cantonese 3.12%). The model focuses on improving general recognition accuracy, contextual intent understanding, professional scene adaptation (supporting hotword injection), and stability in complex acoustic environments (high noise, whispers). The model is now available via API on Tencent Cloud and is integrated into the Yuanbao app.
Related event: Tencent Hunyuan Releases Next-Gen HyASR3.0 Speech Recognition Model(2 posts)→
More from Models
- July Recap: AI Gets Cheap as Open-Source Catches Up — Dapper-Tale-4021 · 2026-08-04
- Swapping Codex to Kimi K3: 1/4 the Cost of GPT-5.6 for Marketing Plans, Minor Quality Drop — AskItAll_Biych · 2026-08-04
- Tencent Hunyuan Launches HyASR3.0 Speech Model, Reducing Multilingual WER to ~3% — 机器之心 · 2026-08-04
- ChatGPT is Breaking Codebases: Devs Complain About Model Regression — wowa93 · 2026-08-04
- DeepSeek V4 Costs 1% of Claude: China's AI Price War Disrupts the Market — SirBoboGargle · 2026-08-04
- OpenAI Hints Next-Gen Models Need More Compute, Codex May Shift to Cloud Agents — haider1 · 2026-08-04