AIToday
Audio & SpeechITmedia AI+Published: Sep 2, 2026, 16:00 JST2 min read

Meta Releases Muse Voice Transcribe, a Real-Time Speech Recognition Model

Meta Releases Muse Voice Transcribe, a Real-Time Speech Recognition Model

Key takeaway

  • Meta launched Muse Voice Transcribe, a real-time speech recognition model.

  • It handles over 20 speakers and audio longer than an hour.

  • The model ranks top in streaming speech recognition and speaker separation benchmarks, with a 3.1% error rate.

3 Key Points

  1. What happened

    Meta has announced Muse Voice Transcribe, its first real-time speech recognition model developed by Meta Superintelligence Labs (MSL). It is available from September 1 through the Meta Model API at $0.18 per hour.

  2. Why it matters

    The model can distinguish more than 20 speakers and handle audio over an hour without post-processing, using a single model for streaming recognition, speaker separation, and end-pointing. It was trained on 70+ languages, with 25 verified at launch, including Japanese, and natively handles code-switching.

  3. What to watch

    Meta claims it ranks #1 in Artificial Analysis' streaming speech recognition and public speaker separation benchmarks. Its word error rate is 3.1% for finalized transcripts, with a 0.16-second delay from end of speech, beating Cartesia Ink-2 (3.4%) and ElevenLabs' Scribe v2 Realtime (3.6%). The model is also integrated into Mac's Meta AI and Muse Code for voice input.

Ask the AI about this article →

Context & Analysis

Muse Voice Transcribe marks Meta's entry into real-time speech recognition with an all-in-one model that combines tasks traditionally handled by separate systems. The technical design uses 80-millisecond audio chunks and adaptive delay learning to balance accuracy and latency, which is a practical approach for real-time applications. Meta's positioning against competitors like Cartesia and ElevenLabs suggests a focus on quality, as shown by the lower error rates. However, the model is not open-sourced; it is available only via API and Meta's own apps, which could lead to different adoption paths for developers. The integration into Mac's Meta AI and Muse Code indicates Meta's intention to embed this technology into its products. As Meta improves the model, it may expand its language support or release open weights, but no such plans are stated.

FAQ

What is the price for using Muse Voice Transcribe via API?
The price on Meta Model API is $0.18 per hour.
Which languages are supported at launch?
The model was trained on over 70 languages, and 25 are verified at the initial release, including Japanese.
Can I use Muse Voice Transcribe in my app?
Yes, it is available to developers through Meta Model API, and also integrated into Mac's Meta AI and Muse Code for voice input.

Get the latest Audio & Speech news every morning

For example, today's edition would include:

  • Winamp Group's Jamendo expands AI music lawsuits, adds six more targetsYahoo Finance AI · 2h ago
  • Google、Gemini 3.7 Flash公開、Pixel 11発表Google AI Blog · 12h ago
  • Phonely launches Alma, voice AI trained on 10M callsSiliconANGLE AI · 17h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleWinamp Group's Jamendo expands AI music lawsuits, adds six more targets