Microsoft released MAI-Transcribe-1, MAI-Voice-1, and MAI-Image-2 on Wednesday, covering speech-to-text, voice generation, and image creation.
The models are available immediately through Microsoft Foundry and a new MAI Playground, targeting enterprise AI applications.
The launch represents Microsoft's first major output from its superintelligence team, formed six months ago by Mustafa Suleyman to achieve 'AI self-sufficiency.'
The three modalities—transcription, voice, and image generation—represent the most commercially valuable areas in enterprise AI.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Aranya Inc., a startup founded last year, launched today with $11 million in funding
CBTS Technology Solutions LLC launched Forge Agents, a platform that turns a plain-language job description in…
Phonely Ltd. launched Alma, a large language AI model built for voice agents and trained on over 10 million re…
Imec CEO Patrick Vandenameele said at SEMICON Taiwan 2026 that the Belgian research center is broadening its c…

Alphabet's AI Overviews now reach over 2.5 billion monthly users through Google Search, and its ad business ge…

Sarah O’Connor's book 'We Are Not Machines' explores how mechanization and AI have transformed the workforce…
