Tencent Hunyuan open-sources AuK, a foundational model unifying speech generation and editing

Tencent-Hunyuan · hf · 2026-09-09

Tencent Hunyuan released AuK, an open-source foundational model that unifies speech generation and editing via natural-language instructions and audio context.

Technically, AuK builds on a multimodal language model with a joint VAE, a hybrid rectified-flow Transformer, and efficient distillation for fast inference. The weights are open-sourced, positioning it as a base model for TTS and speech-editing tasks.

Original post →

More from Multimodal

Multimodal channel →