ByteDance Launches SeedRealtime Full-Duplex Model on Doubao App
aigclink · x · 2026-08-05
ByteDance has released SeedRealtime, a native audio-video full-duplex large model that natively integrates audio, video, and text modalities. It processes continuous video streams directly to see, listen, and speak simultaneously.
Unlike GPT-Live's delegated reasoning approach, SeedRealtime incorporates sound, visuals, timing, and expression into a single model for real-time decision-making. Perception, understanding, decision-making, and expression are synchronized, allowing auditory and visual inputs to jointly participate in real-time judgments. This mechanism enables the model to resolve ambiguous references (like "this") by combining the current visual scene, gestures, gaze, and historical actions, while also disambiguating homophones or unclear speech using visual context.
The model is now available on the Doubao App. Users can experience it by initiating a video call within the latest version of the app.
Related event: ByteDance Launches SeedRealtime Full-Duplex Model on Doubao(9 posts)→
More from Multimodal
- MiniMax Video Workflow Tip: Tweaking Reference Node Settings for Better Face Likeness — xDFINx · 2026-08-06
- Testing Hailuo: Generating Emotional Video with Regional Accents — Disastrous_Coast7870 · 2026-08-06
- Seeking Midjourney Alternatives That Prioritize Creativity Over Photorealism — Vozka · 2026-08-06
- Midjourney Test: Using SREF Code to Create Tense, Dark Supernatural Art — tisch_eins · 2026-08-06
- ComfyUI Video Generation Hurdle: Forced Speech Censorship Ruins Output — YouNoTypey · 2026-08-06
- Suno Unveils Responsible AI Music Principles and Transparency Tools — suno · 2026-08-06