AuK: an open-source foundational model unifying speech generation and editing via natural language
_akhaliq · x · 2026-09-10
The AuK technical report introduces an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. The paper is available on Hugging Face, with the model open-sourced.
More from Multimodal
- Programmable World Model Separates Explicit State Evolution From Video Generation — Zheng-Hui Huang · 2026-09-10
- ByteDance's AgenticGen Uses Business Feedback and Human Rewards to Guide Ad Video Generation — ByteDance · 2026-09-10
- Flash-BoN: cheap drafts beat guided search for diffusion inference-time scaling, +8% AUC at scale — gowthami_s · 2026-09-10
- Alibaba open-sources 2 more speech enhancement models, completing a 7-model toolkit with 22ms echo cancellation — aigclink · 2026-09-10
- A 78-second AI food comedy: Pig Girl crashes a Viking winter feast — Substantial_Depth744 · 2026-09-10
- Shareable ChatGPT Prompt Turns Photos into Mecha-Style Art — Promptmethus · 2026-09-10