ByteDance Launches SeedRealtime Full-Duplex Model on Doubao App

aigclink · x · 2026-08-05

ByteDance has released SeedRealtime, a native audio-video full-duplex large model that natively integrates audio, video, and text modalities. It processes continuous video streams directly to see, listen, and speak simultaneously.

Unlike GPT-Live's delegated reasoning approach, SeedRealtime incorporates sound, visuals, timing, and expression into a single model for real-time decision-making. Perception, understanding, decision-making, and expression are synchronized, allowing auditory and visual inputs to jointly participate in real-time judgments. This mechanism enables the model to resolve ambiguous references (like "this") by combining the current visual scene, gestures, gaze, and historical actions, while also disambiguating homophones or unclear speech using visual context.

The model is now available on the Doubao App. Users can experience it by initiating a video call within the latest version of the app.

Related event: ByteDance Launches SeedRealtime Full-Duplex Model on Doubao(9 posts)→

Original post →

More from Multimodal

Multimodal channel →