ByteDance Releases SeedAudio 1.0 and Opens API

ByteDance's Seed team released the audio creation model SeedAudio 1.0 and announced that its API service is now officially available via Volcano Engine. This model elevates audio generation from producing isolated elements to unified creation of complete soundscapes, allowing developers direct access for integration.

Core Capabilities and Key Details

SeedAudio 1.0 jointly models human voices, sound effects, and ambient sounds within a unified framework. For inputs, the model supports a combination of text prompts and reference audio. For generation control, it features 100ms-level temporal precision, allowing users to set the total track length and control the duration of each line. Additionally, the model can generate up to 2 minutes of audio per attempt and supports generating expressive speech in 20 languages from a single prompt. According to hands-on tests shared by Zhidx, the model can generate narrative-driven audio content relying solely on scene descriptions.

Product Positioning and Impact

Unlike traditional audio generation tools, BytePlus emphasized the "director-level" control experience of SeedAudio 1.0. Users are no longer just passively generating sounds; they can meticulously shape audio details. This combination of precise timing control and multilingual generation significantly lowers the barrier for orchestrating multilingual audio content.

2026-07-20 ~ 2026-07-21 · 5 related posts