Tencent open-sources AuK, a unified 1.5B speech generation and editing model

aigclink · x · 2026-09-11

Tencent has open-sourced AuK, a speech foundation model unifying cloning, instruction-based synthesis, editing, enhancement and separation in one model. It standardizes data as "one instruction + input audio + target voice", letting a single model handle all speech generation and editing tasks. Capabilities include converting normal speech to whispers, removing accents, adding laughter/sighs, and generating voices from scratch. A 4.5x faster lightweight variant targets speed-sensitive use cases like voice assistants, customer service and batch dubbing. Demo: one-person multi-role AI audio drama pipeline.

Related event: Tencent's Open-Source Speech Model AuK Tops Hugging Face Trending(2 posts)→

Original post →

More from Multimodal

Multimodal channel →