
Cohere has released Cohere Transcribe Arabic, an open-source speech-to-text model designed specifically for Arabic's linguistic challenges.
The 2-billion-parameter system outperforms existing models like Whisper Large V3 on human-rated benchmarks for quality, dialect handling, and code-switching accuracy.
It is freely available under an open license, making it accessible to developers and organizations working with Arabic speech.
What happened
Cohere released Cohere Transcribe Arabic, a 2-billion-parameter open-source model for Arabic speech recognition. The model is available on Hugging Face and through the Cohere API under the Apache 2.0 license.
Why it matters
According to Cohere, it is the most accurate open-source Arabic speech-to-text system available and outperforms Whisper Large V3 and the standard Cohere Transcribe model in benchmarks. It addresses Arabic's specific challenges—dialect variety, bilingual Arabic-English conversations, code-switching, and specialized vocabulary—which are difficult for general speech recognition systems to handle accurately.
What to watch
Human ratings on a 1–5 scale show Cohere Transcribe Arabic outperforms both Whisper Large V3 and the standard Cohere Transcribe model in overall quality, dialect faithfulness, and code-switching. The model is available now on Hugging Face and via the Cohere API.
Ask the AI about this article →
Cohere's release of Cohere Transcribe Arabic reflects a targeted approach to a persistent gap in open-source speech recognition: Arabic's linguistic complexity has often been underserved by general-purpose models. The model's 2-billion-parameter size and open-source licensing lower barriers to adoption for researchers and developers working across the Arabic-speaking world, removing the dependency on proprietary or English-optimized systems. By benchmarking against both Whisper Large V3 (a widely-used baseline) and its own standard transcription model, Cohere anchors the improvement to concrete reference points, suggesting the gains come from Arabic-specific training and design rather than a general capability jump. The emphasis on code-switching and dialect faithfulness points to real-world Arabic speech environments where mixing languages and regional variation are the norm rather than exceptions.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
AI system scaling has pushed interconnect requirements inside data centers from chips and boards up to racks…

Chinese large-model developer Z.ai says it can now support large-scale inference using roughly 100,000 domesti…

Analyst Ming-Chi Kuo says Nvidia has revived the Rubin CPX AI accelerator with a substantially redesigned arch…

Palantir Technologies stock has posted multi-year gains, including an 11x return over 3 years

Apple has escalated its legal battle against OpenAI, claiming in a new court filing that OpenAI is actively de…

Samsung Electronics has locked up as much as 70% of its memory production capacity under long-term supply agre…
