Tencent Hunyuan Launches HyASR3.0 Speech Model, Reducing Multilingual WER to ~3%

机器之心 · wechat · 2026-08-04

Tencent Hunyuan has released HyASR3.0 preview, its latest speech recognition model. Built on the Hy3 LLM, the model adopts a MoE architecture and a self-developed unsupervised speech encoder to fuse high-precision ASR with deep semantic understanding.

On open-source benchmarks, HyASR3.0 preview achieves a Word Error Rate (WER) of around 3% across multiple languages (3.34% for Mandarin, 2.62% for English, and 3.12% for Cantonese). It introduces significant improvements in general recognition accuracy, context-aware smart error correction, hotword injection for professional scenarios, and stability in complex acoustic environments like high noise or whispers. The model is now available via Tencent Cloud APIs and integrated into the Yuanbao app.

Related event: Tencent Hunyuan Releases Next-Gen HyASR3.0 Speech Recognition Model(2 posts)→

Original post →

More from Models

Models channel →