
Meta launched Muse Voice Transcribe, a real-time speech recognition model.
It handles over 20 speakers and audio longer than an hour.
The model ranks top in streaming speech recognition and speaker separation benchmarks, with a 3.1% error rate.
What happened
Meta has announced Muse Voice Transcribe, its first real-time speech recognition model developed by Meta Superintelligence Labs (MSL). It is available from September 1 through the Meta Model API at $0.18 per hour.
Why it matters
The model can distinguish more than 20 speakers and handle audio over an hour without post-processing, using a single model for streaming recognition, speaker separation, and end-pointing. It was trained on 70+ languages, with 25 verified at launch, including Japanese, and natively handles code-switching.
What to watch
Meta claims it ranks #1 in Artificial Analysis' streaming speech recognition and public speaker separation benchmarks. Its word error rate is 3.1% for finalized transcripts, with a 0.16-second delay from end of speech, beating Cartesia Ink-2 (3.4%) and ElevenLabs' Scribe v2 Realtime (3.6%). The model is also integrated into Mac's Meta AI and Muse Code for voice input.
Ask the AI about this article →
Muse Voice Transcribe marks Meta's entry into real-time speech recognition with an all-in-one model that combines tasks traditionally handled by separate systems. The technical design uses 80-millisecond audio chunks and adaptive delay learning to balance accuracy and latency, which is a practical approach for real-time applications. Meta's positioning against competitors like Cartesia and ElevenLabs suggests a focus on quality, as shown by the lower error rates. However, the model is not open-sourced; it is available only via API and Meta's own apps, which could lead to different adoption paths for developers. The integration into Mac's Meta AI and Muse Code indicates Meta's intention to embed this technology into its products. As Meta improves the model, it may expand its language support or release open weights, but no such plans are stated.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Winamp Group's subsidiary Jamendo SA amended its U.S

Google released Gemini 3.7 Flash, its latest developer model, and unveiled the new Pixel 11 series

Phonely Ltd. launched Alma, a large language AI model built for voice agents and trained on over 10 million re…
Two developers released TontaubeV1, a 2.9B-parameter open-weight TTS model for expressive speech and long-form…

Mitsubishi Electric and its U.S

EDM producer Max "H4RRIS" Harris and Italian turntablist-turned-producer Nihil Young are publicly calling out…
