ByteDance’s SwanTale targets multi-speaker speech and audio generation
ByteDance · hf · 2026-08-04
ByteDance unveils SwanTale, a multi-speaker speech and audio generation model
The paper targets two core scenarios: instruct generation, where users describe the environment, speaker style, and content in natural language, and zero-shot generation, where reference audio is used together with the same fine-grained content.
- ByteDance also introduces SwanData-Caption, a new data pipeline that cleans raw speech/audio, adds synthetic coverage, and produces multi-level captions.
- On the model side, SwanTale adds SwanVAE, reward-conditioned quality control, Engram conditioning, and a Unified MoE design.
- The authors say the system is trained with curriculum learning and GRPO post-training.
- Reported results show SwanTale leads on multiple zero-shot and instruct metrics, with particularly strong expressiveness and support for complex multi-speaker generation.
A demo site is linked in the post.
More from Multimodal
- MiniMax H3 prompt turns lesson ideas into a 15-second cinematic typography video — umesh_ai · 2026-08-04
- Alibaba’s UEmbed unifies sparse and dense multimodal retrieval in one model — Alibaba-NLP · 2026-08-04
- WorldExam benchmarks whether video models generate worlds that actually react — Yuxue Yang · 2026-08-04
- MiniMax H3 users ask whether reference video input can take a loaded video node — Nilakshexplains · 2026-08-04
- Seedance 2.5 gets mocked as “job-worthy” a year ago, scroll-past today — andrew_n_carr · 2026-08-04
- ReToken adds one learned embedding to improve long-context visual retrieval — burkov · 2026-08-04