Source: MarkTechPost
Microsoft AI has released MAI-Transcribe-2-Streaming, its first streaming speech-to-text (STT) model. It launched on October 1, 2026, alongside 2 text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Artificial Analysis ranks it #1 of 38 models for final and first partial transcript accuracy. The model targets voice agents, live captions and dictation, where latency decides the experience.
What Microsoft Shipped
MAI-Transcribe-2-Streaming is the real-time sibling of the batch MAI-Transcribe-2, released in September. It transcribes 60 languages with automatic, continuous language detection. Audio streams in continuously, and text streams back while the speaker is still talking.
The model emits its first hypotheses, called partials, just over 100ms after receiving audio. It revises those partials as context arrives, then commits a stable final transcript. An agent can therefore start reasoning or calling tools mid-sentence. Microsoft team states its internal tests show words appearing 2x faster than its closest competitor.
What Artificial Analysis Measured
The AA-WER Streaming index uses about 8 hours of audio. The mix is AA-AgentTalk (50%), VoxPopuli (25%) and Earnings22 (25%). Latency is timed from the end of speech, as detected by SileroVAD.
- Final transcript: 2.5% WER at 0.13s after end of speech, #1 of 38 models.
- First partial transcript: 2.5% WER at 0.12s after end of speech, also #1.
- Runners-up: Grok Voice Transcribe 2.0 at 2.7% and 0.49s. Muse Voice Transcribe at 3.1% and 0.16s.
- Not the fastest: Cartesia Ink-2 (external endpoints) returns finals in 0.07s, but at 4.0% WER.
The first partial is as accurate as the final transcript. That matters for agents that act before the speaker finishes. Microsoft also places the model on the accuracy versus latency Pareto frontier.
Pricing
MAI-Transcribe-2-Streaming costs $0.54 per hour of audio. This is an introductory price through the end of 2026. Artificial Analysis normalizes it to $9.00 per 1,000 minutes. Batch MAI-Transcribe-2 costs $0.10 per hour. On streaming, Microsoft charges more than xAI and Meta, and roughly matches Google’s estimated rate.
How Developers Integrate It
Microsoft documents 2 integration paths. The Realtime API suits apps already using an OpenAI Realtime-compatible WebSocket. The Azure Speech SDK handles connection management, retries and audio streaming. Both return intermediate and final results.
The model is also available in the MAI Playground, through Vercel and Azure Voice Live. LiveKit support is listed as coming soon.
Microsoft pairs it with MAI-Voice-2.1-Flash for full voice loops. Flash generates 45s of audio at 150ms end-to-end latency for $15 per 1M characters. MAI-Voice-2.1 covers 23 languages and 26 locales at $22 per 1M characters.
Comparison: MAI-Transcribe-2-Streaming vs Closest Streaming Competitors
| Feature | MAI-Transcribe-2-Streaming | Grok Voice Transcribe 2.0 | Muse Voice Transcribe | Gemini 3.5 Transcribe Live |
|---|---|---|---|---|
| Developer | Microsoft AI | xAI | Meta Superintelligence Labs | |
| Released | Oct 1, 2026 | Sep 18, 2026 | Sep 1, 2026 | Aug 26, 2026 |
| AA-WER Streaming (final) | 2.5% | 2.7% | 3.1% | 4.0% |
| Time to final | 0.13s | 0.49s | 0.16s | Not reported by AA source cited |
| Streaming price / hour | $0.54 (intro) | $0.20 | $0.18 | ~$0.54 (token-billed estimate) |
| Languages | 60, continuous auto-detect | Dozens, auto-detect, mid-recording switch | 70+ trained, 25 verified | 85+, auto-detect |
| Speaker diarization in stream | Not stated | Included in API (streaming not confirmed) | Yes, 20+ speakers | Not supported in Live mode |
| Interface | Realtime API (WebSocket) + Azure Speech SDK | WebSocket | WebSocket + file endpoint | Live API (WebSocket) |
| Open weights | No | No | No | No |
Sources: Artificial Analysis, 9to5Mac (Muse pricing), DataCamp (Grok pricing), The Batch (Gemini pricing). Verified October 2, 2026.
Key Takeaways
- #1 of 38 on AA-WER Streaming: 2.5% WER at 0.13s to final.
- First partials score the same 2.5% WER, at 0.12s.
- 60 languages with continuous automatic language detection.
- $0.54 per hour introductory price, higher than xAI and Meta.
- Public preview with no SLA and no open weights.
Check out the Technical details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to partner with us for promoting your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar etc.? Connect with us
Michal Sutter
Michal Sutter is a data science professional with a Master of Science in Data Science from the University of Padova. With a solid foundation in statistical analysis, machine learning, and data engineering, Michal excels at transforming complex datasets into actionable insights.

