AIToday
Large Language ModelsAudio & SpeechTHE DECODERPublished: Oct 2, 2026, 19:01 JST

Microsoft AI's MAI-Transcribe-2-Streaming tops accuracy, 100ms latency

Microsoft AI's MAI-Transcribe-2-Streaming tops accuracy, 100ms latency

3 Key Points

  1. What happened

    Microsoft AI launched MAI-Transcribe-2-Streaming, which Microsoft says ranks first for accuracy on Artificial Analysis, handles 60 languages, and returns first partial results in just over 100 milliseconds. An hour of audio costs $0.54 at the introductory price through the end of the year.

  2. Why it matters

    The sub-100-millisecond partial results are designed to let voice agents respond while someone is still mid-sentence, which could make automated phone and chat agents feel more like a live conversation.

  3. What to watch

    Whether low latency and voice cloning hold up in real deployments, since both voice models can clone a voice from a few seconds of audio and Microsoft says built-in safeguards are meant to prevent misuse — in one test about half of 4,000 participants thought the voices belonged to a real person.

WHO IT HITSCompanies and developers building voice agents — customer-support phone systems, voice assistants, call-center automation — can use the lower latency and pricing; teams handling sensitive calls have to weigh the voice-cloning and misuse safeguards Microsoft says are built in.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Microsoft AI is pairing a transcription engine with two text-to-speech models in a single voice-agent lineup. The transcription model's headline claim is speed and accuracy: Microsoft says it ranks first for accuracy on Artificial Analysis, covers 60 languages, and returns first partial results in just over 100 milliseconds, which Microsoft says lets voice agents respond while someone is still mid-sentence. Pricing is an introductory $0.54 per hour of audio through the end of the year.

The two speech models extend the same push. MAI-Voice-2.1 is designed to speak 23 languages in the same voice with a native accent in each, and the MAI-Voice-2.1-Flash variant is listed at 150 milliseconds of latency and $15 per million characters instead of $22. Both voice models can clone a voice from a few seconds of reference audio, and Microsoft says built-in safeguards are meant to prevent misuse. Distribution runs through Microsoft Foundry and the MAI Playground, and the two voice models are also on OpenRouter.

The adoption question hinges on trust as much as latency. In one test, about half of 4,000 participants thought the voices belonged to a real person — a result that cuts both ways for products that depend on voice agents sounding human, and for the safeguards Microsoft says are meant to keep cloning from being misused. How regulators and buyers react to that trade-off is likely to shape how quickly the models move from playground demos into live customer-facing systems.

FAQ
Where can developers use the new Microsoft voice models?
They are available through Microsoft Foundry and the MAI Playground, among other platforms, and the two voice models are also on OpenRouter.
How much does the transcription model cost?
Through the end of the year, an hour of audio costs $0.54 at the introductory price.
What languages do the new models support?
MAI-Transcribe-2-Streaming transcribes 60 languages, and MAI-Voice-2.1 is designed to speak 23 languages in the same voice with a native accent in each one.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleELYZA founds ELYZA RSI Research for AI-driven AI development