
What happened
Microsoft AI said it released MAI-Transcribe-2-Streaming, which returns provisional results in just over 100 milliseconds and ranked first in accuracy in an Artificial Analysis evaluation. It supports 60 languages and costs $0.54 per hour of audio as introductory pricing.
Why it matters
A model that returns results in just over 100 milliseconds lets voice agents begin processing while the other person is still speaking, which Microsoft views as the basis for the voice agents it is targeting.
What to watch
The accuracy ranking comes from Artificial Analysis's streaming transcription evaluation, and the $0.54 hourly price is introductory pricing through the end of the year, so the standing rate is likely to change.
WHO IT HITSThis lands on developers and product teams building voice agents, who gain a low-latency transcription option priced per hour of audio, and on enterprises that need voice cloning to require access approval and the speaker's consent.
Summaries like this, in your inbox every morning.
Microsoft AI, the AI division of Microsoft, announced three models together, and the range of languages involved is telling. The streaming transcription model handles 60 languages including Japanese, while the multilingual text-to-speech model covers 23 languages, and the voice models can keep one voice while switching languages with native accents. The pricing is set per use: $0.54 per hour of audio, $22 per million characters, and $15 per million characters for the high-speed Flash version, which Microsoft says can generate 45 seconds of audio with 150 milliseconds of latency.
The release places Microsoft against other large players the article names as already moving into voice AI, including OpenAI's streaming transcription model, Google's live model, and Meta's open-source speech recognition system. Microsoft's answer is to bundle transcription and speech synthesis for building voice agents, and to offer them not only in Microsoft Foundry but also in MAI Playground and elsewhere, alongside a combined demo called Chatter. In that demo, speaking Japanese is recognized, but responses are not configured in Japanese.
The voice replication feature is the part with the clearest guardrails: Microsoft says it requires review and approval, and only voices with the person's consent can be synthesized. What all of this ultimately hinges on is how builders respond to the accuracy ranking Microsoft cites from Artificial Analysis and to the low-latency claims, balanced against the introductory pricing that runs only through the end of the year.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Microsoft launched MAI-Transcribe-2-Streaming, its first streaming transcription model, priced at 54 cents per…
Google announced Gemini 4 Argon on September 30, saying DeepMind's own evaluation beat GPT-6 Astra, Claude Fab…

Microsoft released MAI-Transcribe-2-Streaming on October 1, 2026

Anthropic said it made the web service claude.ai and its desktop app about 3 times faster in 2 weeks, and that…

On October 1, OpenAI updated ChatGPT's release notes with shopping features — a 'try on' button on product car…

At the Lytham Partners Fall 2026 Investor Conference, Rezolve AI CEO Dan Wagner said the company generated $13…
