Tencent releases AuK, a unified speech generation and editing model with a 4.5x faster Flash version

_akhaliq · x · 2026-09-10

Tencent released AuK, a unified speech generation and editing model that handles zero-shot TTS, instruction-driven content/acoustic/paralinguistic editing, plus enhancement and separation through one natural-language interface.

The distilled AuK-Flash variant runs 4.5x faster.

Related event: Tencent Hunyuan open-sources unified speech model AuK(3 posts)→

Original post →

More from Multimodal

Multimodal channel →