ByteDance Releases SeedAudio 1.0 and Opens API

ByteDance Seed released the audio creation model SeedAudio 1.0, and Volcano Engine officially opened its API service for developers to directly integrate. The model pushes audio generation capabilities from single-element production to unified creation of complete sound scenes, attracting significant attention.

Core Capabilities and Key Details

SeedAudio 1.0 jointly models human voices, sound effects, and ambient sounds within a unified framework. For input, the model supports a combination of text prompts and reference audio. Regarding generation control, it features 100ms-level time precision, allowing users to set the overall track length and control the duration of each line. Additionally, the model can generate approximately 2 minutes of audio per session and supports expressive speech generation in 20 languages using a single prompt. According to hands-on testing shared by Zhidx, the model can generate dramatic audio content relying solely on scene descriptions.

Product Positioning and Impact

Unlike traditional audio generation tools, BytePlus emphasized the "director-level" control experience of SeedAudio 1.0 in its introduction. Users are no longer just passively generating sound; they can finely shape audio details. This combination of precise time control and multilingual generation significantly lowers the barrier for multilingual audio content arrangement.

2026-07-20 ~ 2026-07-21 · 5 related posts

Primary sources