Microsoft Launches MAI-Transcribe-2-Streaming and Two MAI-Voice Models

microsoft-launches-mai-transcribe-2-streaming-and-two-mai-voice-models

Source: Unite.AI

Microsoft AI on October 1, 2026 launched MAI-Transcribe-2-Streaming, its first streaming transcription model, alongside two new text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, with all three available through Microsoft Foundry.

MAI-Transcribe-2-Streaming and the Artificial Analysis Benchmark

Microsoft describes MAI-Transcribe-2-Streaming as delivering low-latency, real-time transcripts in 60 languages with automatic, continuous language detection. The company said the model ranks No. 1 for accuracy for both final and partial transcripts on Artificial Analysis, and that it sits on the Pareto frontier of the benchmark’s accuracy-versus-latency evaluation, meaning higher accuracy does not require a heavy latency tradeoff. The leaderboard chart in the post, citing the Artificial Analysis streaming leaderboard dated September 28, 2026, shows the model at a 2.5 percent final word-error rate, a 2.8 percent first-partial rate, and 0.13 seconds to final transcription.

Rather than waiting for a speaker to finish before returning text, the model produces its first hypotheses, known as partials, in just over 100 milliseconds of receiving audio, then revises them as more context arrives before committing a stable transcript. Microsoft said this allows voice-enabled applications to act on speech before the speaker finishes: voice agents can start reasoning or calling tools mid-sentence, and live transcripts can appear as people talk. For real-time dictation and subtitling, the company said its internal evaluations show words appearing in the transcript twice as fast as with its closest competitor.

Artificial Analysis states that its AA-WER Streaming index measures transcription accuracy for models where audio streams in real time, chunk by chunk, across roughly eight hours of audio from three datasets: AA-AgentTalk at 50 percent, VoxPopuli at 25 percent, and Earnings22 at 25 percent. The datasets cover real-world speech with diverse accents, domain-specific language, and challenging acoustic conditions, and the benchmark’s Time to Final and Time to First Partial measurements both start at the end of speech detected by the SileroVAD voice-activity detector.

MAI-Transcribe-2-Streaming is available at an introductory price of $0.54 per hour of audio through the end of the year. The model extends Microsoft’s MAI audio line, which already includes MAI-Transcribe-2, the earlier non-streaming speech recognition model the company billed as the fastest, most accurate and cheapest in the world.

MAI-Voice-2.1 and MAI-Voice-2.1-Flash

MAI-Voice-2.1 supports 23 languages and 26 locales, and Microsoft said a single voice can use all of them with a native accent, keeping the same speaker identity when switching languages. A tutoring app, in the company’s example, can switch languages mid-lesson without swapping teachers, and a multilingual assistant can reply in whatever language it is addressed in while still sounding like the same voice. The model is priced at $22 per 1M characters.

MAI-Voice-2.1-Flash supports the same languages and cross-language speakers but is built for high-volume, latency-sensitive workloads. It can generate up to 45 seconds of audio with an end-to-end latency of 150 milliseconds, and Microsoft said it delivers 55 percent faster model inference and is roughly 60 percent cheaper than comparable models. It is priced at $15 per 1M characters.

Both voice models support voice cloning across all supported languages using a few seconds of reference audio, with built-in consent guardrails that Microsoft said prevent misuse. In a 4,000-listener Turing test combining the two new voice models, 50.3 percent of listeners rated MAI-Voice as equally or more human-like than human recordings, Microsoft said.

Microsoft framed pairing MAI-Transcribe-2-Streaming with MAI-Voice-2.1-Flash as buying back time on both ends of a voice-agent loop, the sequence of hearing, understanding, deciding, and speaking within the window where a human still experiences the interaction as a conversation. Listed developer use cases include customer service agents that transcribe requests as they are spoken and respond in natural speech, multilingual assistants that detect the spoken language and reply in any of the 23 supported MAI-Voice languages, and interactive learning and media applications using distinct speakers for tutoring, role-play, simulations, narration, and conversational content.

Availability and the Chatter Demo

MAI-Voice-2.1 and MAI-Voice-2.1-Flash are available through OpenRouter. All three models are available through Microsoft Foundry, the MAI Playground, Vercel, and Azure Voice Live, with LiveKit listed as coming soon.

To show the models working together in a live agent, Microsoft built Chatter, a new demo in the MAI Playground that lets users talk to a voice assistant powered by the transcription and voice models.

mm

Jonas Reeve is an AI-generated analyst at Unite.AI, focusing on cognitive AI, artificial general intelligence (AGI), and the theoretical foundations of machine intelligence. His work explores how learning, reasoning, memory, and abstraction emerge in both biological and artificial systems, drawing connections between modern AI architectures and long-standing questions in cognitive science and philosophy of mind.

With a conceptual and reflective approach, Jonas examines frameworks such as reasoning models, agentic systems, emergent cognition, and alignment theory, aiming to clarify what progress toward AGI actually means—and what it does not. Rather than chasing timelines or hype, he emphasizes first principles, conceptual rigor, and the limits of current models.

Articles authored by Jonas Reeve are AI-generated and reviewed by Unite.AI’s editorial team to ensure accuracy, clarity, and responsible discussion of advanced AI concepts.