AIToday
Large Language ModelsAudio & SpeechITmedia AI+Published: Oct 2, 2026, 13:00 JST

Microsoft AI unveils first streaming transcription model MAI-Transcribe-2-Streaming

Microsoft AI unveils first streaming transcription model MAI-Transcribe-2-Streaming

3 Key Points

  1. What happened

    Microsoft AI said it released MAI-Transcribe-2-Streaming, which returns provisional results in just over 100 milliseconds and ranked first in accuracy in an Artificial Analysis evaluation. It supports 60 languages and costs $0.54 per hour of audio as introductory pricing.

  2. Why it matters

    A model that returns results in just over 100 milliseconds lets voice agents begin processing while the other person is still speaking, which Microsoft views as the basis for the voice agents it is targeting.

  3. What to watch

    The accuracy ranking comes from Artificial Analysis's streaming transcription evaluation, and the $0.54 hourly price is introductory pricing through the end of the year, so the standing rate is likely to change.

WHO IT HITSThis lands on developers and product teams building voice agents, who gain a low-latency transcription option priced per hour of audio, and on enterprises that need voice cloning to require access approval and the speaker's consent.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Microsoft AI, the AI division of Microsoft, announced three models together, and the range of languages involved is telling. The streaming transcription model handles 60 languages including Japanese, while the multilingual text-to-speech model covers 23 languages, and the voice models can keep one voice while switching languages with native accents. The pricing is set per use: $0.54 per hour of audio, $22 per million characters, and $15 per million characters for the high-speed Flash version, which Microsoft says can generate 45 seconds of audio with 150 milliseconds of latency.

The release places Microsoft against other large players the article names as already moving into voice AI, including OpenAI's streaming transcription model, Google's live model, and Meta's open-source speech recognition system. Microsoft's answer is to bundle transcription and speech synthesis for building voice agents, and to offer them not only in Microsoft Foundry but also in MAI Playground and elsewhere, alongside a combined demo called Chatter. In that demo, speaking Japanese is recognized, but responses are not configured in Japanese.

The voice replication feature is the part with the clearest guardrails: Microsoft says it requires review and approval, and only voices with the person's consent can be synthesized. What all of this ultimately hinges on is how builders respond to the accuracy ranking Microsoft cites from Artificial Analysis and to the low-latency claims, balanced against the introductory pricing that runs only through the end of the year.

FAQ
How much does Microsoft's streaming transcription model cost?
MAI-Transcribe-2-Streaming is priced at $0.54 per hour of audio as introductory pricing through the end of the year. The normal rate is not stated.
Can Microsoft's new voice models clone someone's voice?
Yes, both Voice models can replicate a voice from a few seconds of reference audio. Use requires a review process and access approval, and the system only synthesizes voices for which the person's consent has been obtained.
Is Japanese supported by Microsoft's new voice models?
Japanese is not yet in the default voice list published in the announcement. However, Japan East is one of the regions where the models are offered, and the streaming transcription model supports Japanese among its 60 languages.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleNvidia at 14.6x next year's earnings: top pick for October