Velocity
This Week's Stories
ByteDance

ByteDance

AI · Multimodal AI · Consumer Apps

Launch

ByteDance launches SeedRealtime full-duplex audio-video AI model

August 5, 2026

The model adds live visual understanding on top of voice, cutting interruption and error rates roughly in half versus stitched-together systems.

  • ByteDance officially launched SeedRealtime, a native audio-video full-duplex large model, on August 5, 2026, now fully rolled out in the Doubao app.
  • SeedRealtime is ByteDance's latest Seed-family model, built for real-time multimodal interaction rather than as a standalone chatbot.
  • The model uses a unified architecture to natively fuse audio, video, and text, enabling a continuous 'see, listen, and speak simultaneously' interaction stream instead of separate perception and response pipelines.
  • Users can point a camera at objects, follow along with a process, or read text aloud, and the system proactively flags changes or mistakes it observes in real time.
  • In noisy, multi-speaker settings, SeedRealtime combines speaker identification with face recognition, voice, and gesture understanding, cutting interruptions, lag, and noise-triggered errors by about half compared with prior integrated systems.
  • SeedRealtime succeeds ByteDance's earlier Seeduplex system, which handled only voice and text, marking the addition of real-time vision as the next step toward full 'omni-modal' interaction.
  • The launch shows ByteDance racing to match and surpass Western rivals like Google's Gemini in live multimodal AI, embedding the tech directly into its billion-user Doubao consumer app rather than a developer-only tool.

Read More About This Story

Get the app

Stay Ahead With Velocity

Deep company profiles, investor context, and every original source behind this story — plus the next one, the moment it breaks.

Download on the App StoreGet it on Google Play

More This Week

View All →