Doubao Launches Omni-Modal Full-Duplex Model: Is It Better Than GPT Live?
vista8 · x · 2026-08-06
ByteDance's Doubao App has fully rolled out a new video call feature powered by SeedRealtime, a native audio-visual full-duplex large model. It integrates sound, vision, timing, and expression into a single model for real-time decision-making, achieving full-duplex interaction across video, audio, and text.
Technical Breakthroughs & Experience
- Eliminating Cascading Latency: Unlike traditional multi-module cascades (ASR, VLM, TTS) or half-duplex systems relying on external VADs, SeedRealtime processes perception, understanding, decision-making, and expression simultaneously.
- Visual Disambiguation: When users use deictic words like "this," the model combines the visual scene, user gestures, and gaze direction to identify the specific reference.
- Real-World Tests: In testing, the model successfully identified fleeting figures in a YouTube video and interpreted charts on screen. Compared to OpenAI's audio-only full-duplex GPT Live, SeedRealtime offers lower latency, more natural Chinese voice output without an accent, and true visual understanding.
The feature is now available to all users in the latest Doubao App version by clicking the "call" icon and enabling the camera.
Related event: ByteDance Launches SeedRealtime Full-Duplex Model on Doubao(9 posts)→
More from Models
- Dev Seeks Western-Hosted Platforms for GLM and Kimi Models — sgt102 · 2026-08-06
- DeepSeek API Surprises Developers with Built-in Web Search Integration — max_paperclips · 2026-08-06
- MiniMax-H3-experimental Tops Hugging Face Trending Models — Kijai · 2026-08-06
- Alibaba's Qwen3.8 Max: 2.4T MoE Matches Claude Opus, Weights Coming Next Week — ArtificialAnlys · 2026-08-06
- Alexandr Wang Amazed: Muse Spark 1.2 Shows Unprecedented Capabilities — alexandr_wang · 2026-08-06
- MiniMax H3 Tops Video Generation Categories on DesignArena — airesearch12 · 2026-08-06