AuK: an open-source foundational model unifying speech generation and editing via natural language

_akhaliq · x · 2026-09-10

The AuK technical report introduces an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. The paper is available on Hugging Face, with the model open-sourced.

Related event: Tencent Hunyuan Open-Sources AuK, a Unified Speech Generation and Editing Model(4 posts)→

Original post →

More from Multimodal

Multimodal channel →