
What happened
Microsoft AI launched MAI-Transcribe-2-Streaming, which Microsoft says ranks first for accuracy on Artificial Analysis, handles 60 languages, and returns first partial results in just over 100 milliseconds. An hour of audio costs $0.54 at the introductory price through the end of the year.
Why it matters
The sub-100-millisecond partial results are designed to let voice agents respond while someone is still mid-sentence, which could make automated phone and chat agents feel more like a live conversation.
What to watch
Whether low latency and voice cloning hold up in real deployments, since both voice models can clone a voice from a few seconds of audio and Microsoft says built-in safeguards are meant to prevent misuse — in one test about half of 4,000 participants thought the voices belonged to a real person.
WHO IT HITSCompanies and developers building voice agents — customer-support phone systems, voice assistants, call-center automation — can use the lower latency and pricing; teams handling sensitive calls have to weigh the voice-cloning and misuse safeguards Microsoft says are built in.
Summaries like this, in your inbox every morning.
Microsoft AI is pairing a transcription engine with two text-to-speech models in a single voice-agent lineup. The transcription model's headline claim is speed and accuracy: Microsoft says it ranks first for accuracy on Artificial Analysis, covers 60 languages, and returns first partial results in just over 100 milliseconds, which Microsoft says lets voice agents respond while someone is still mid-sentence. Pricing is an introductory $0.54 per hour of audio through the end of the year.
The two speech models extend the same push. MAI-Voice-2.1 is designed to speak 23 languages in the same voice with a native accent in each, and the MAI-Voice-2.1-Flash variant is listed at 150 milliseconds of latency and $15 per million characters instead of $22. Both voice models can clone a voice from a few seconds of reference audio, and Microsoft says built-in safeguards are meant to prevent misuse. Distribution runs through Microsoft Foundry and the MAI Playground, and the two voice models are also on OpenRouter.
The adoption question hinges on trust as much as latency. In one test, about half of 4,000 participants thought the voices belonged to a real person — a result that cuts both ways for products that depend on voice agents sounding human, and for the safeguards Microsoft says are meant to keep cloning from being misused. How regulators and buyers react to that trade-off is likely to shape how quickly the models move from playground demos into live customer-facing systems.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Ramp economist Ara Kharazian says US firms are spending less on AI even as usage rose about 50 percent from Ju…

ELYZA said it is launching "ELYZA RSI Research", and that for LLMs of 100 billion parameters or fewer it has r…

A Fortune commentator writes that in the July OpenAI sandbox incident, agents that attacked Hugging Face left…

Superhuman, the productivity company formerly known as Grammarly, agreed in June to acquire GPTZero for undisc…

Earendil released Pi 1.0 on October 1, 2026, with standard MCP support, Codemode, Deferred tool loading, virtu…

Researchers at the Objective-See Foundation found a now-patched bug in the macOS ChatGPT app that could take o…
