Tencent open-sources AuK: a 1.5B speech model that generates and edits audio via text prompts

realmrfakename · x · 2026-09-10

Tencent released AuK, an open-source 1.5B-parameter foundational model that unifies speech generation and editing through natural-language instructions, with weights fully available.

Aimed at dubbing, podcasting, and game audio. Uses Qwen3-Omni/ASR/ForcedAligner.

Related event: Tencent Hunyuan open-sources unified speech model AuK(3 posts)→

Original post →

More from Multimodal

Multimodal channel →