AIToday

OpenAI's GPT Transcribe improves 0.7 points but trails ElevenLabs, Google, Mistral

THE DECODER12h agoSend on LINE
OpenAI's GPT Transcribe improves 0.7 points but trails ElevenLabs, Google, Mistral

Key takeaway

OpenAI has launched GPT Transcribe and GPT Live Transcribe, new speech recognition models that achieve a 3.31 percent word error rate—a 0.7 percentage point improvement over the prior version—while cutting pricing 25 percent to $0.0045 per minute. However, the models rank fourth on the AA-WER benchmark behind ElevenLabs, Google, and Mistral, facing direct price competition from Mistral's offering at $0.003 per minute.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    OpenAI released GPT Transcribe and GPT Live Transcribe, two speech recognition models via its API. GPT Transcribe processes pre-recorded audio about 34 times faster than real time, while GPT Live Transcribe handles real-time streaming with low latency. The new GPT Transcribe achieves a 3.31 percent word error rate—a 0.7 percentage point improvement over GPT-4o Transcribe—at a 25 percent price cut to $0.0045 per minute of audio.

  • Why it matters

    OpenAI's transcription models now rank #4 on the AA-WER benchmark, behind ElevenLabs Scribe v2 (2.3%), Google's Gemini 3 Pro (2.9%), and Mistral's Voxtral Small (3%). The price drop positions OpenAI competitively against Mistral's Voxtral Transcribe V2, which starts at $0.003 per minute—suggesting OpenAI is responding to margin pressure in a crowded speech recognition market.

  • What to watch

    The models accept text as transcription context, keywords, and multiple input languages, and they complement OpenAI's recently announced Realtime model generation, which includes GPT-Realtime-Whisper. Full technical details are available in OpenAI's Transcription Guide.

In Depth

OpenAI has introduced two new speech recognition models through its API: GPT Transcribe, designed for pre-recorded audio, and GPT Live Transcribe, built for real-time streaming. GPT Transcribe processes audio at approximately 34 times faster than real time, making it efficient for batch workloads. GPT Live Transcribe prioritizes low-latency real-time performance for live applications.

According to the AA-WER benchmark maintained by Artificial Analysis, GPT Transcribe achieves a word error rate of 3.31 percent. This represents a 0.7 percentage point improvement compared to GPT-4o Transcribe, its one-year-old predecessor. Concurrent with the model release, OpenAI reduced pricing by 25 percent, bringing the cost to $0.0045 per minute of audio. Both models accept text as transcription context, keywords, and support multiple input languages.

In the competitive landscape, OpenAI's transcription models rank fourth on the AA-WER benchmark. ElevenLabs Scribe v2 leads with a 2.3 percent error rate, followed by Google's Gemini 3 Pro at 2.9 percent and Mistral's Voxtral Small at 3 percent. Mistral has recently pressured the market by introducing Voxtral Transcribe V2 at $0.003 per minute, undercutting OpenAI's new pricing. These new transcription models complement OpenAI's recently announced Realtime model generation, which also includes the real-time transcription model GPT-Realtime-Whisper. Full technical details are available in OpenAI's Transcription Guide.

Context & Analysis

OpenAI's release of GPT Transcribe and GPT Live Transcribe represents a incremental but meaningful improvement in its transcription capabilities. The 0.7 percentage point reduction in word error rate and 25 percent price drop signal an effort to maintain competitiveness in a market where ElevenLabs, Google, and Mistral have established stronger benchmarks. ElevenLabs Scribe v2 leads the AA-WER ranking at 2.3 percent, while Google's Gemini 3 Pro (2.9%) and Mistral's Voxtral Small (3%) both outperform OpenAI's offering. More notably, Mistral's recent move to $0.003 per minute—underselling OpenAI's $0.0045 price point—appears to have prompted OpenAI's own pricing adjustment. The models' support for text context, keywords, and multiple input languages positions them as flexible tools, and their integration with OpenAI's recently announced Realtime model generation (including GPT-Realtime-Whisper) suggests the company is bundling transcription with broader real-time AI capabilities. Whether the combination of modest accuracy gains and aggressive pricing can close the gap with faster competitors remains to be seen.

FAQ

What is the word error rate for GPT Transcribe?
GPT Transcribe achieves a 3.31 percent word error rate according to the AA-WER benchmark, a 0.7 percentage point improvement over GPT-4o Transcribe.
How does GPT Transcribe's pricing compare to competitors?
GPT Transcribe costs $0.0045 per minute of audio after a 25 percent price cut. Mistral's Voxtral Transcribe V2 starts at $0.003 per minute, undercutting the market.
What is the difference between GPT Transcribe and GPT Live Transcribe?
GPT Transcribe processes pre-recorded audio files about 34 times faster than real time, while GPT Live Transcribe is built for real-time streaming with low latency.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime