
What happened
Meta's Superintelligence Labs released Muse Voice Transcribe, a real-time audio perception model that transcribes speech, distinguishes speakers, and detects sentence boundaries during live conversation. It is available through the Meta Model API and powers voice dictation in Meta AI and Muse Code.
Why it matters
The model undercuts rivals on price at $0.18 per hour ($3 per 1,000 audio minutes), while achieving a 3.1% word error rate on English in independent testing, compared to 3.6% for ElevenLabs Scribe v2 Realtime and 4.0% for AssemblyAI. It can handle over 20 speakers and recordings over an hour long without post-processing.
What to watch
The model's success hinges on whether its adaptive delay mechanism—which balances accuracy against speed—proves reliable in real-world use. Watch for how competitors like OpenAI, which released GPT-Realtime-Whisper in May, respond on price.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
Meta's release of Muse Voice Transcribe marks its entry into the real-time speech recognition arena, following its reorganization under Superintelligence Labs in mid-2025. The model's adaptive delay mechanism—which decides per word how long to listen—represents a technical approach aimed at balancing accuracy and speed, trained via reinforcement learning to optimize both. This positions Meta to compete on price, a strategy consistent with its previous Muse Spark releases, rather than on raw performance alone.
The independent evaluation by Artificial Analysis, showing a 3.1% error rate, suggests Meta's model is competitive with established players like ElevenLabs and AssemblyAI, while undercutting them on cost. This could pressure rivals to adjust pricing, especially as OpenAI has already cut transcription prices in July. However, Meta's decision not to release weights or disclose training details may limit adoption among developers who prefer open models.
The stakes are tied to Zuckerberg's vision of "personal superintelligence," where reliable speech recognition underpins AI agents that listen through glasses. Yet regulatory hurdles, such as the recent discussion in Germany about banning Meta's camera glasses, could slow this vision. The model's success may hinge on whether its price advantage and accuracy convince developers to choose it over alternatives, despite the closed-source approach.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Google has released its Lyria 3.5 music generation model in the Gemini app and via API

AWS published a solution to deploy a multimodal WhatsApp ordering assistant using Amazon Bedrock AgentCore and…

Tokyu Construction announced on August 31, 2026, that it will use NTT ConoSurf's voice AI and generative AI to…

Roland introduced Melody Flip, a plug-in for digital audio workstations that generates musical ideas

Microsoft AI released MAI-Transcribe-2, a speech-recognition model it claims is faster, more accurate, and che…

The article explores how Vocaloid, Yamaha's singing-synthesizer software, has been used by Japanese artists to…
