Tencent Hunyuan Launches HyASR3.0 Speech Model, Reducing Multilingual WER to ~3%
机器之心 · wechat · 2026-08-04
Tencent Hunyuan has released HyASR3.0 preview, its latest speech recognition model. Built on the Hy3 LLM, the model adopts a MoE architecture and a self-developed unsupervised speech encoder to fuse high-precision ASR with deep semantic understanding.
On open-source benchmarks, HyASR3.0 preview achieves a Word Error Rate (WER) of around 3% across multiple languages (3.34% for Mandarin, 2.62% for English, and 3.12% for Cantonese). It introduces significant improvements in general recognition accuracy, context-aware smart error correction, hotword injection for professional scenarios, and stability in complex acoustic environments like high noise or whispers. The model is now available via Tencent Cloud APIs and integrated into the Yuanbao app.
Related event: Tencent Hunyuan Releases Next-Gen HyASR3.0 Speech Recognition Model(2 posts)→
More from Models
- Testing GPT-5.6 and Gemini Robotics Models: Impressive but Need Real Deployment Data — m_wulfmeier · 2026-08-04
- Far Behind GPT? Devs Complain Gemini Live Lacks Emotional Nuance — TheBuzzer4625kHz · 2026-08-04
- July Recap: AI Gets Cheap as Open-Source Catches Up — Dapper-Tale-4021 · 2026-08-04
- Swapping Codex to Kimi K3: 1/4 the Cost of GPT-5.6 for Marketing Plans, Minor Quality Drop — AskItAll_Biych · 2026-08-04
- DeepSeek V4 Costs 1% of Claude: China's AI Price War Disrupts the Market — SirBoboGargle · 2026-08-04
- OpenAI Hints Next-Gen Models Need More Compute, Codex May Shift to Cloud Agents — haider1 · 2026-08-04