Microsoft ships streaming transcription and two new voice models
Microsoft's AI unit closed out Wednesday with three speech models at once — a streaming transcriber and two voices — aimed squarely at the workloads where latency is the product: live captioning, voice agents, and the contact-center stack.
Microsoft released MAI-Transcribe-2-Streaming, its first real-time transcription model, alongside two voice models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. The transcriber covers 60 languages with automatic language detection between them, and it starts returning partial transcripts in just over 100 milliseconds of receiving audio — before the speaker has finished the sentence. All three ship through Microsoft Foundry, the MAI Playground and Azure Voice Live, with Vercel also carrying them and OpenRouter listing the voice models. Microsoft is putting the streaming model at $0.54 an audio hour, an introductory price good through the end of 2026.
The claim to watch is the accuracy-at-speed one: on Artificial Analysis's streaming leaderboard, Microsoft's model leads both measures — a 2.5% word error rate on final transcripts and 2.8% on first partials, with the final version landing 0.13 seconds after the speaker stops. That combination, sub-second finalization with single-digit error, is the threshold at which live transcription stops feeling like captions catching up and starts feeling like simultaneous interpretation. On the voice side, MAI-Voice-2.1 spans 23 languages and 26 locales while holding the same voice identity across all of them, at $22 per million characters; Flash is the latency play at 150 milliseconds end to end, which Microsoft says is 55% faster inference and roughly 60% cheaper than comparable models, at $15 per million characters. Microsoft also ran a listening test with 4,000 participants in which its generated voice was rated at human quality or better 50.3% of the time.
Why it matters: speech is currently the most price-competitive corner of the model market, and Microsoft is competing on third-party-checked numbers rather than demo reels — product lead Mustafa Suleyman framed the new voice model directly against ElevenLabs as "55% faster and 60% cheaper." The distribution tells as much of the story as the benchmarks: landing on Vercel and OpenRouter means MAI models can be picked up outside Azure's walled garden, the same route Qwen and DeepSeek used to get into other companies' apps. It picks up the real-time audio race we tracked last month — Meta ships Muse Voice Transcribe, its first real-time audio model — and it raises the bar for all of them: Microsoft is publishing error rates at millisecond latencies, which makes "fast" a number competitors now have to beat rather than a word they get to use.
What to watch: whether the $0.54 introductory rate survives into 2027, and how fast ElevenLabs and the hyperscalers answer on streaming accuracy.
If the streaming numbers hold up in your own pipeline, is switching transcription vendors still a six-month project, or does a third off the price make it a Tuesday? Tell us in the comments.