Microsoft Expands MAI Family With Streaming Transcription and Voice Models
Microsoft has introduced three new MAI artificial intelligence models, featuring its first streaming transcription model alongside two text-to-speech variants for voice agents.

Microsoft Corp. has expanded its MAI artificial intelligence model family with the introduction of its first streaming transcription model, alongside two new text-to-speech models. These additions are built for developers seeking to construct voice agents capable of listening to human speech and replying instantly, mirroring natural human conversation.
The most notable release in this rollout is MAI-Transcribe-2-Streaming. According to the company, this model accepts human speech via a WebSocket and transforms it into a transcript that updates continuously as the person speaks. Once the user finishes talking, the model confirms that the transcript is final. This functionality allows applications to display live captions or start processing a user request before the person has even finished speaking.
MAI-Transcribe-2-Streaming is available on Microsoft’s Vercel AI Gateway at a price of $0.54 per audio hour. It supports over 60 languages and automatically detects which one is being spoken. Microsoft states that it delivers its first transcript hypotheses within 320 milliseconds on average, though actual speed depends on network conditions and the AI system generating the response.
This streaming option follows last month's debut of the regular MAI-Transcribe-2 model, which is priced at $0.10 per audio hour—more than five times cheaper than the streaming version. The price difference stems from the heavy processing required by the streaming model to continuously return provisional results while audio arrives, whereas the standard model waits for the speaker to finish.
Voice Generation and Overall Architecture
Beyond transcription, Microsoft released two models focused on text-to-speech generation. MAI-Voice-2.1 is designed to deliver expressive and high-fidelity outputs, while MAI-Voice-2.1-Flash lowers costs and trades off some of those attributes for faster response times. Listed on the Vercel AI Gateway, MAI-Voice-2.1 costs $22 per million characters, and the Flash variant is priced at $15 per million characters. Both text-to-speech models support 23 languages.
Building a functional voice AI agent requires three steps: recognizing speech, reasoning about the input, and generating an audible response. MAI-Transcribe-2-Streaming handles the initial recognition, while the MAI-Voice models manage the final audio generation. Sitting in the middle is Microsoft’s reasoning model, Mai-Thinking-1, a large language model that processes transcripts to determine the agent's action. This multi-model approach gives developers granular control over cost, latency, and quality.
What it means for developers
These new models reflect Microsoft's ongoing push to build out its proprietary MAI ecosystem, driven by AI Chief Executive Mustafa Suleyman's goal to reduce and eventually eliminate the costs associated with external frontier models from providers like OpenAI and Anthropic. For developers wanting to build responsive conversational applications, having separate components for transcription, reasoning, and speech synthesis offers a modular framework. Additionally, developers can try top AI models cheaply through one API at https://apixoai.online.
The rollout provides everything necessary to build voice agents that interact naturally with users, pointing toward a future where Microsoft's own Copilot agents across Excel, Outlook, and other platforms could be powered entirely by the in-house MAI family.
Source: Microsoft targets ultra-realistic voice agents with its first streaming transcription model — SiliconANGLE AI. Written by the Apixo team from that report.
One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.
Get your API key

