OpenAI introduced three audio models in the API: GPT-Realtime-2 (a voice model with GPT-5-class reasoning), GPT-Realtime-Translate (supporting 70+ input languages and 13 output languages for live translation), and GPT-Realtime-Whisper (a streaming speech-to-text for live transcription).
GPT-Realtime-2 increases context window from 32K to 128K to support longer sessions, adds adjustable reasoning effort (minimal, low, medium, high, xhigh), and includes preambles and parallel tool calls that make agent actions audible to users during task completion.
GPT-Realtime-2 (high) scores 15.2% higher on Big Bench Audio for audio intelligence than GPT-Realtime-1.5; GPT-Realtime-2 (xhigh) scores 13.8% higher on Audio MultiChallenge for instruction following. Zillow reported a 26-point lift in call success rate after prompt optimization (95% vs. 69%) on adversarial benchmarks.
Ask the AI about this article →
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Goldman Sachs, Morgan Stanley and Citigroup are pressing elite law firms to lower fees, arguing AI is sharply…

MediaTek shares closed 10% higher on Tuesday after the Taiwanese chip firm announced a partnership with Nvidia

Dell raised its annual revenue outlook, citing strong sales of AI servers

GE Vernova (NYSE: GEV) announced on August 24, 2026, in Paris, a new medium-voltage uninterruptible power supp…

Oracle (ORCL) stock is down 4% to $142.82 in Tuesday afternoon trading, after the 10-year Treasury yield climb…

Tesla's investments in AI startups lag far behind its Magnificent Seven peers
