Tencent open-sources AuK: a 1.5B speech model that generates and edits audio via text prompts
realmrfakename · x · 2026-09-10
Tencent released AuK, an open-source 1.5B-parameter foundational model that unifies speech generation and editing through natural-language instructions, with weights fully available.
- Five task families: speech generation, content editing, enhancement/separation, paralinguistic editing (swap accents, inject emotion), and acoustic editing (remove noise, replace a single spoken word)
- Zero-shot voice cloning without reference transcripts, plus full TTS, vocal separation, and studio-grade cleanup
- Massive training scale: 3.03 billion instruction-audio instances and 1.95 million hours of effective supervision; reported 2.65% error rate
- Architecture: multimodal LLM for semantic conditioning + audio VAE trained on speech, general audio, and music + hybrid rectified-flow Transformer with dual-stream MMDiT blocks
- AuK-Flash: distilled via trajectory-level consistency init and task-routing decoupled DMD, reaching near-teacher quality in 4 steps without CFG
Aimed at dubbing, podcasting, and game audio. Uses Qwen3-Omni/ASR/ForcedAligner.
Related event: Tencent Hunyuan open-sources unified speech model AuK(3 posts)→
More from Multimodal
- Fable 5.1 Demos 'Walking Inside a Van Gogh Painting', Feat No Other Model Nails — thursdai_pod · 2026-09-10
- Open-source YuE2 music model generates editable symbolic scores before rendering full songs — GreyScope · 2026-09-10
- Redditor re-renders Robot Chicken sketch in classic SpongeBob art style with Wan and Minimax — Ok-Giraffe-8670 · 2026-09-10
- AI-Generated 1990s-Style Sitcom 'Andrew & Kale' Airs Weekly Episodes — letandrewcook · 2026-09-10
- MiniMax H3 image generation wows users in a chess 'checkmate' test run — charis_ai · 2026-09-10
- This prompt turns any AI image into a candid behind-the-scenes iPhone shot — flowersslop · 2026-09-10