Tencent open-sources AuK: a 1.5B speech model unifying generation and instruction-based editing
pmttyji · reddit · 2026-09-12
Tencent's Hunyuan team open-sourced AuK (paper: arXiv 2609.08936), a 1.5B foundation model for speech generation and editing trained on millions of hours of audio, with every task exposed through a unified natural-language instruction interface:
- Generation: zero-shot TTS from reference audio, instruct TTS from a voice description alone;
- Content editing: rewriting what is said, lyric editing that preserves melody and voice;
- Acoustic editing: pitch (semitones), speed, volume;
- Paralinguistic editing: emotion, timbre, accent removal, adding/removing breaths/laughs/coughs, whisper conversion;
- Enhancement & separation: denoising, speech enhancement, source separation.
Two variants: AuK (base) and AuK-Flash (distilled, fast 4-step inference), weights on HF and ModelScope, plus a Cookbook with instruction templates and CLI/Python examples.
Related event: Tencent Hunyuan Open-Sources AuK All-in-One Speech Model(4 posts)→
More from Multimodal
- Blur-to-video: SIGGRAPH Asia 2025 work recovers past, present and future frames from one motion-blurred photo — CSProfKGD · 2026-09-12
- Using Blender camera references with Seeddance 2.5 to produce a polished AI ad video — aziz4ai · 2026-09-12
- AI-Generated 'Celestial Goddess Performs Every 400 Years' Video Goes Viral — Less-Daikon-250 · 2026-09-12
- Overhead hoop, mid-wave hand — AI nails physical consistency in a group photo — misovalko · 2026-09-12
- GPT-Images 2.5 shows strong grasp of products and brand guidelines — EXM7777 · 2026-09-12
- One prompt drives an entire chase: Seedance 2.5 demo on RunwayML — azed_ai · 2026-09-12