
What happened
Meta has announced Muse Voice Transcribe, its first real-time speech recognition model developed by Meta Superintelligence Labs (MSL). It is available from September 1 through the Meta Model API at $0.18 per hour.
Why it matters
The model can distinguish more than 20 speakers and handle audio over an hour without post-processing, using a single model for streaming recognition, speaker separation, and end-pointing. It was trained on 70+ languages, with 25 verified at launch, including Japanese, and natively handles code-switching.
What to watch
Meta claims it ranks #1 in Artificial Analysis' streaming speech recognition and public speaker separation benchmarks. Its word error rate is 3.1% for finalized transcripts, with a 0.16-second delay from end of speech, beating Cartesia Ink-2 (3.4%) and ElevenLabs' Scribe v2 Realtime (3.6%). The model is also integrated into Mac's Meta AI and Muse Code for voice input.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
Muse Voice Transcribe marks Meta's entry into real-time speech recognition with an all-in-one model that combines tasks traditionally handled by separate systems. The technical design uses 80-millisecond audio chunks and adaptive delay learning to balance accuracy and latency, which is a practical approach for real-time applications. Meta's positioning against competitors like Cartesia and ElevenLabs suggests a focus on quality, as shown by the lower error rates. However, the model is not open-sourced; it is available only via API and Meta's own apps, which could lead to different adoption paths for developers. The integration into Mac's Meta AI and Muse Code indicates Meta's intention to embed this technology into its products. As Meta improves the model, it may expand its language support or release open weights, but no such plans are stated.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Google has released its Lyria 3.5 music generation model in the Gemini app and via API

Meta's Superintelligence Labs released Muse Voice Transcribe, a real-time audio perception model that transcri…

AWS published a solution to deploy a multimodal WhatsApp ordering assistant using Amazon Bedrock AgentCore and…

Tokyu Construction announced on August 31, 2026, that it will use NTT ConoSurf's voice AI and generative AI to…

Roland introduced Melody Flip, a plug-in for digital audio workstations that generates musical ideas

Microsoft AI released MAI-Transcribe-2, a speech-recognition model it claims is faster, more accurate, and che…
