ByteDance Launches SeedRealtime, a Native Full-Duplex Audio-Video Model on Doubao

字节跳动Seed · wechat · 2026-08-05

ByteDance's Seed team has officially released SeedRealtime, a native full-duplex audio-video large model. Using a unified end-to-end architecture that natively integrates audio, video, and text, it breaks through the high latency and information loss bottlenecks of traditional cascade systems, enabling real-time "see, hear, and speak" interactions.

Core capabilities include:

End-to-end human evaluations show a 50% reduction in conversational rhythm issues compared to cascade models. SeedRealtime is now fully rolled out in the Doubao App, accessible via the video call feature.

Original post →

More from Multimodal

Multimodal channel →