Microsoft Unveils MAI-Transcribe-2-Streaming Model
Microsoft AI has launched MAI-Transcribe-2-Streaming, ranking #1 on Artificial Analysis for real-time speech-to-text accuracy and speed.

Microsoft AI has officially introduced MAI-Transcribe-2-Streaming, marking the organization's entry into real-time speech-to-text systems. Debuting on October 1, 2026, the streaming variant arrived alongside two text-to-speech offerings: MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Designed specifically for live captions, voice agents, and dictation environments, the model focuses heavily on minimizing latency where speed dictates user experience. Functioning as the real-time counterpart to the batch-oriented MAI-Transcribe-2 released in September, the new system processes 60 languages equipped with continuous, automatic language detection.
Performance and Benchmarks
Artificial Analysis evaluated the system using the AA-WER Streaming index, which comprises approximately 8 hours of audio balanced across AA-AgentTalk, VoxPopuli, and Earnings22 datasets. Latency measurements began immediately at the end of speech, identified via SileroVAD. MAI-Transcribe-2-Streaming achieved the top position out of 38 models for both final and first partial transcript accuracy.
The model returns its initial hypotheses, termed partials, just over 100 milliseconds after receiving audio inputs. It continuously revises these partials as additional context arrives before committing to a stable final transcript. Benchmarks recorded a 2.5% Word Error Rate at 0.13 seconds following the end of speech for the final transcript, and a matching 2.5% WER at 0.12 seconds for the first partial transcript. Competitors included Grok Voice Transcribe 2.0 at 2.7% WER and 0.49 seconds, and Muse Voice Transcribe at 3.1% WER and 0.16 seconds. While external endpoints like Cartesia Ink-2 returned finals faster at 0.07 seconds, they registered a higher 4.0% WER.
Pricing and Integration
Introductory pricing runs at $0.54 per hour of audio through the end of 2026, which normalizes to $9.00 per 1,000 minutes. This rate positions Microsoft higher than streaming offerings from xAI and Meta, while roughly matching estimated rates from Google. Batch MAI-Transcribe-2 remains available at $0.10 per hour.
Developers can access the technology through two primary integration pathways. The Realtime API accommodates applications already utilizing an OpenAI Realtime-compatible WebSocket, while the Azure Speech SDK manages connections, retries, and audio streaming. Both routes supply intermediate and final results. Furthermore, the model is accessible via the MAI Playground, Azure Voice Live, and Vercel, with LiveKit integration noted as coming soon. Developers looking to experiment with various top AI models cheaply through one API can explore options at https://apixoai.online.
What it means for developers
For engineers building conversational voice agents, MAI-Transcribe-2-Streaming introduces capabilities that allow systems to reason or trigger tool calls mid-sentence. Because the first partial transcript achieves the same 2.5% accuracy level as the final text, applications can safely act on user input before a speaker finishes talking. The availability of standardized integration endpoints—such as the OpenAI Realtime-compatible WebSocket and the Azure Speech SDK—simplifies the implementation of real-time audio streams into existing application architectures.
Source: Microsoft AI Releases MAI-Transcribe-2-Streaming: #1 Real-Time Speech-to-Text Model on Artificial Analysis — MarkTechPost. Written by the Apixo team from that report.
One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.
Get your API key

