AIToday
Audio & SpeechAI Business & IndustryOpenAI BlogPublished: May 8, 2026, 04:00 JST1 min read

OpenAI launches GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper audio models for developers building voice applications.

3 Key Points

  1. OpenAI introduced three audio models in the API: GPT-Realtime-2 (a voice model with GPT-5-class reasoning), GPT-Realtime-Translate (supporting 70+ input languages and 13 output languages for live translation), and GPT-Realtime-Whisper (a streaming speech-to-text for live transcription).

  2. GPT-Realtime-2 increases context window from 32K to 128K to support longer sessions, adds adjustable reasoning effort (minimal, low, medium, high, xhigh), and includes preambles and parallel tool calls that make agent actions audible to users during task completion.

  3. GPT-Realtime-2 (high) scores 15.2% higher on Big Bench Audio for audio intelligence than GPT-Realtime-1.5; GPT-Realtime-2 (xhigh) scores 13.8% higher on Audio MultiChallenge for instruction following. Zillow reported a 26-point lift in call success rate after prompt optimization (95% vs. 69%) on adversarial benchmarks.

Ask the AI about this article →

Get the latest Audio & Speech news every morning

For example, today's edition would include:

  • Phonely launches Alma, voice AI trained on 10M callsSiliconANGLE AI · 6h ago
  • Mitsubishi Electric develops task-general sound separation AITop Companies AI · 1d ago
  • Musician Detectives Hunt AI Music GriftersThe Verge AI · 3d ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleCommerce.com reports 5% revenue growth and returns to GAAP profitability in Q1 2026 as AI-driven commerce initiatives expand