ByteDance Launches SeedRealtime Full-Duplex Audio-Visual LLM

testingcatalog · x · 2026-08-05

ByteDance's Seed team has officially launched SeedRealtime, a native audio-visual full-duplex large language model. Built on a unified architecture, the model processes continuous audio, video, and text streams simultaneously, enabling real-time interaction where it can watch, listen, and speak concurrently.

The core breakthrough is the model's ability to autonomously determine conversational timing without relying on external Voice Activity Detection (VAD) rules. It tracks scenes, identifies speakers, and understands pauses and background noise. By integrating visual context, it resolves linguistic ambiguities like homophones. In open-environment tests such as noisy restaurants or group dinners, the model demonstrated highly accurate face/voice matching and object recognition, even proactively issuing reminders when a requested item appeared on screen.

Related event: ByteDance Launches Full-Duplex SeedRealtime Model on Doubao(5 posts)→

Original post →

More from Models

Models channel →