
ByteDance
AI · Multimodal AI · Consumer Apps
ByteDance launches SeedRealtime full-duplex audio-video AI model
August 5, 2026
The model adds live visual understanding on top of voice, cutting interruption and error rates roughly in half versus stitched-together systems.
- ByteDance officially launched SeedRealtime, a native audio-video full-duplex large model, on August 5, 2026, now fully rolled out in the Doubao app.
- SeedRealtime is ByteDance's latest Seed-family model, built for real-time multimodal interaction rather than as a standalone chatbot.
- The model uses a unified architecture to natively fuse audio, video, and text, enabling a continuous 'see, listen, and speak simultaneously' interaction stream instead of separate perception and response pipelines.
- Users can point a camera at objects, follow along with a process, or read text aloud, and the system proactively flags changes or mistakes it observes in real time.
- In noisy, multi-speaker settings, SeedRealtime combines speaker identification with face recognition, voice, and gesture understanding, cutting interruptions, lag, and noise-triggered errors by about half compared with prior integrated systems.
- SeedRealtime succeeds ByteDance's earlier Seeduplex system, which handled only voice and text, marking the addition of real-time vision as the next step toward full 'omni-modal' interaction.
- The launch shows ByteDance racing to match and surpass Western rivals like Google's Gemini in live multimodal AI, embedding the tech directly into its billion-user Doubao consumer app rather than a developer-only tool.