ByteDance Releases SeedAudio1.0

字节跳动Seed · wechat · 2026-07-20

ByteDance Seed has released SeedAudio1.0, an audio creation model designed to generate complete soundscapes rather than isolated audio elements.

The model jointly models vocals, sound effects, and ambient sounds within a unified framework. It supports text + reference audio inputs, 100ms-level time control, generates roughly 2 minutes of audio per run with extension capabilities, and produces natural audio in 20+ languages. The official release also included evaluations across film, podcasts, and livestream e-commerce scenarios, claiming a usability rate of over 90% for most cases.

Related event: ByteDance Releases SeedAudio 1.0 and Opens API(5 posts)→

Original post →

More from Multimodal

Multimodal channel →