Tencent Hunyuan open-sources AuK, a foundational model unifying speech generation and editing
Tencent-Hunyuan · hf · 2026-09-09
Tencent Hunyuan released AuK, an open-source foundational model that unifies speech generation and editing via natural-language instructions and audio context.
Technically, AuK builds on a multimodal language model with a joint VAE, a hybrid rectified-flow Transformer, and efficient distillation for fast inference. The weights are open-sourced, positioning it as a base model for TTS and speech-editing tasks.
More from Multimodal
- DaVinci Resolve Can Now Use Claude/ChatGPT as Your Personal Editing Assistant — _AustinCalvert_ · 2026-09-09
- AI video editing agents shift AI from generation to workflow execution — OwlZealousideal4779 · 2026-09-09
- AI-Generated Video of a Cat Riding a Roller Coaster at Santa Monica Pier — Born_History_8246 · 2026-09-09
- Video gen is so fast that 'script' may no longer be the right pre-production artifact — tobowers · 2026-09-09
- ComfyUI Style Node NeonsStyleExplorer Adds Custom Catalogs and Style Libraries — neonsparksuk · 2026-09-09
- H3 Character LoRA Training: Likeness Far Behind Wan 2.2, Tips and Pitfalls Shared — Tiny-Highlight-9180 · 2026-09-09