Tencent releases AuK, a unified speech generation and editing model with a 4.5x faster Flash version
_akhaliq · x · 2026-09-10
Tencent released AuK, a unified speech generation and editing model that handles zero-shot TTS, instruction-driven content/acoustic/paralinguistic editing, plus enhancement and separation through one natural-language interface.
The distilled AuK-Flash variant runs 4.5x faster.
Related event: Tencent Hunyuan open-sources unified speech model AuK(3 posts)→
More from Multimodal
- AuK: an open-source foundational model unifying speech generation and editing via natural language — _akhaliq · 2026-09-10
- Same surreal prompt tested: MiniMax H3 passes while Kling and Seedance fail — charis_ai · 2026-09-10
- AI-Generated Elon Musk Covers "Ridin' Dirty" in Eerily Convincing Video — bennash · 2026-09-10
- GPT Image 2.5 goes live on OpenArt with Flare and Sunburst modes — OpenAIDevs · 2026-09-10
- Open-source Minimax Studio unifies local AI video production on ComfyUI — Wild_Ant5693 · 2026-09-10
- Building AI live wallpapers in ComfyUI with MiniMax H3, GIMM-VFI and DLSS 5 — alisitskii · 2026-09-10