Velocity
This Week's Stories
Microsoft

Microsoft

Artificial Intelligence · Voice AI · Speech Recognition

Launch

Microsoft launches MAI-Transcribe-2-Streaming, two new MAI-Voice models

October 1, 2026

The model answers in milliseconds, letting voice agents start reasoning before a caller finishes a sentence.

  • Microsoft launched MAI-Transcribe-2-Streaming alongside two text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, on October 1, 2026, available via Microsoft Foundry and Vercel's AI Gateway.
  • The models target voice-agent builders: Microsoft says the streaming transcriber lets a customer-service agent start identifying a caller's request before the sentence ends, according to Microsoft Foundry product lead Naomi Moneypenny.
  • MAI-Transcribe-2-Streaming is priced at an introductory $0.54 per hour of audio, more than five times the $0.10 rate for Microsoft's non-streaming MAI-Transcribe-2 released a month earlier.
  • On the Artificial Analysis streaming leaderboard, the model posts a 2.5% final word-error rate, the lowest ranked, and returns first partial transcripts within roughly 100 to 320 milliseconds.
  • MAI-Voice-2.1 speaks 23 languages with one consistent voice identity across languages at $22 per million characters, while Flash generates 45 seconds of audio at 150ms latency for $15 per million characters.
  • The launch follows Microsoft's July pledge to spend $2.5 billion on a new 'Microsoft Frontier Company' unit embedding 6,000 engineers with customers to help deploy AI.
  • Pricing streaming transcription at a steep premium over batch transcription shows Microsoft betting developers will pay for lower latency, a key front in the voice-agent infrastructure race against OpenAI, Google and ElevenLabs.

Read More About This Story

Get the app

Stay Ahead With Velocity

Deep company profiles, investor context, and every original source behind this story — plus the next one, the moment it breaks.

Download on the App StoreGet it on Google Play

More This Week

View All →