
Microsoft
Artificial Intelligence · Voice AI · Speech Recognition
Microsoft launches MAI-Transcribe-2-Streaming, two new MAI-Voice models
October 1, 2026
The model answers in milliseconds, letting voice agents start reasoning before a caller finishes a sentence.
- Microsoft launched MAI-Transcribe-2-Streaming alongside two text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, on October 1, 2026, available via Microsoft Foundry and Vercel's AI Gateway.
- The models target voice-agent builders: Microsoft says the streaming transcriber lets a customer-service agent start identifying a caller's request before the sentence ends, according to Microsoft Foundry product lead Naomi Moneypenny.
- MAI-Transcribe-2-Streaming is priced at an introductory $0.54 per hour of audio, more than five times the $0.10 rate for Microsoft's non-streaming MAI-Transcribe-2 released a month earlier.
- On the Artificial Analysis streaming leaderboard, the model posts a 2.5% final word-error rate, the lowest ranked, and returns first partial transcripts within roughly 100 to 320 milliseconds.
- MAI-Voice-2.1 speaks 23 languages with one consistent voice identity across languages at $22 per million characters, while Flash generates 45 seconds of audio at 150ms latency for $15 per million characters.
- The launch follows Microsoft's July pledge to spend $2.5 billion on a new 'Microsoft Frontier Company' unit embedding 6,000 engineers with customers to help deploy AI.
- Pricing streaming transcription at a steep premium over batch transcription shows Microsoft betting developers will pay for lower latency, a key front in the voice-agent infrastructure race against OpenAI, Google and ElevenLabs.