ByteDance Releases SeedAudio 1.0 and Opens API
ByteDance Seed released the audio creation model SeedAudio 1.0, and Volcano Engine officially opened its API service for developers to directly integrate. The model pushes audio generation capabilities from single-element production to unified creation of complete sound scenes, attracting significant attention.
Core Capabilities and Key Details
SeedAudio 1.0 jointly models human voices, sound effects, and ambient sounds within a unified framework. For input, the model supports a combination of text prompts and reference audio. Regarding generation control, it features 100ms-level time precision, allowing users to set the overall track length and control the duration of each line. Additionally, the model can generate approximately 2 minutes of audio per session and supports expressive speech generation in 20 languages using a single prompt. According to hands-on testing shared by Zhidx, the model can generate dramatic audio content relying solely on scene descriptions.
Product Positioning and Impact
Unlike traditional audio generation tools, BytePlus emphasized the "director-level" control experience of SeedAudio 1.0 in its introduction. Users are no longer just passively generating sound; they can finely shape audio details. This combination of precise time control and multilingual generation significantly lowers the barrier for multilingual audio content arrangement.
2026-07-20 ~ 2026-07-21 · 5 related posts
Primary sources
- [source] ByteDance Releases SeedAudio1.0 — 字节跳动Seed · 2026-07-20
- [source] SeedAudio1.0 Opens API with Multilingual Support — 火山引擎 · 2026-07-20
- [source] SeedAudio1.0 generates dialogue, effects, and ambience in one shot — 智东西 · 2026-07-20
- Dola Seed Audio 1.0 gets a capability upgrade — DavidmComfort · 2026-07-20
- BytePlus’s Dola Seed Audio 1.0 adds timing control for 20-language speech generation — iamaliveix · 2026-07-21