AIToday
Audio & SpeechOpen-Source AIAI Business & IndustryTHE DECODERPublished: Jul 8, 2026, 04:01 JST2 min read

Cohere releases open-source Arabic speech-to-text model

Cohere releases open-source Arabic speech-to-text model

Key takeaway

  • Cohere has released Cohere Transcribe Arabic, an open-source speech-to-text model designed specifically for Arabic's linguistic challenges.

  • The 2-billion-parameter system outperforms existing models like Whisper Large V3 on human-rated benchmarks for quality, dialect handling, and code-switching accuracy.

  • It is freely available under an open license, making it accessible to developers and organizations working with Arabic speech.

3 Key Points

  1. What happened

    Cohere released Cohere Transcribe Arabic, a 2-billion-parameter open-source model for Arabic speech recognition. The model is available on Hugging Face and through the Cohere API under the Apache 2.0 license.

  2. Why it matters

    According to Cohere, it is the most accurate open-source Arabic speech-to-text system available and outperforms Whisper Large V3 and the standard Cohere Transcribe model in benchmarks. It addresses Arabic's specific challenges—dialect variety, bilingual Arabic-English conversations, code-switching, and specialized vocabulary—which are difficult for general speech recognition systems to handle accurately.

  3. What to watch

    Human ratings on a 1–5 scale show Cohere Transcribe Arabic outperforms both Whisper Large V3 and the standard Cohere Transcribe model in overall quality, dialect faithfulness, and code-switching. The model is available now on Hugging Face and via the Cohere API.

Ask the AI about this article →

Context & Analysis

Cohere's release of Cohere Transcribe Arabic reflects a targeted approach to a persistent gap in open-source speech recognition: Arabic's linguistic complexity has often been underserved by general-purpose models. The model's 2-billion-parameter size and open-source licensing lower barriers to adoption for researchers and developers working across the Arabic-speaking world, removing the dependency on proprietary or English-optimized systems. By benchmarking against both Whisper Large V3 (a widely-used baseline) and its own standard transcription model, Cohere anchors the improvement to concrete reference points, suggesting the gains come from Arabic-specific training and design rather than a general capability jump. The emphasis on code-switching and dialect faithfulness points to real-world Arabic speech environments where mixing languages and regional variation are the norm rather than exceptions.

FAQ

How does Cohere Transcribe Arabic compare to other systems?
According to Cohere, it is the most accurate open-source Arabic speech-to-text system available and outperforms Whisper Large V3 and the standard Cohere Transcribe model in overall quality, dialect faithfulness, and code-switching.
What languages or use cases does it handle?
The model is built to handle Arabic's specific challenges, including dialect variety, bilingual Arabic-English conversations, code-switching, and specialized vocabulary.
How can I access it?
Cohere Transcribe Arabic is available on Hugging Face and through the Cohere API under the Apache 2.0 license.

Get the latest Audio & Speech news every morning

For example, today's edition would include:

  • Mitsubishi Electric develops task-general sound separation AITop Companies AI · 12h ago
  • Musician Detectives Hunt AI Music GriftersThe Verge AI · 2d ago
  • Beatport bans AI-generated music from DJ marketplaceTHE DECODER · 3d ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articlePalantir CEO attacks OpenAI, Anthropic over IP theft claims—but case mostly weak