Tencent open-sources AuK: a 1.5B speech model unifying generation and instruction-based editing

pmttyji · reddit · 2026-09-12

Tencent's Hunyuan team open-sourced AuK (paper: arXiv 2609.08936), a 1.5B foundation model for speech generation and editing trained on millions of hours of audio, with every task exposed through a unified natural-language instruction interface:

Two variants: AuK (base) and AuK-Flash (distilled, fast 4-step inference), weights on HF and ModelScope, plus a Cookbook with instruction templates and CLI/Python examples.

Related event: Tencent Hunyuan Open-Sources AuK All-in-One Speech Model(4 posts)→

Original post →

More from Multimodal

Multimodal channel →