What happened
Vercel's AI Gateway now supports audio and voice capabilities including real-time voice conversations, text-to-speech, and speech-to-text transcription. These features launch with models from OpenAI and xAI, available in beta through AI SDK 7.
Why it matters
Real-time voice agents let applications hold live conversations where users can interrupt and talk over the model as they would with a person, making it practical for voice assistants, customer support, and hands-free tools. Unlike chaining separate speech-to-text, language model, and text-to-speech steps, a single real-time model hears and produces audio directly.
What to watch
Audio calls route through the same AI Gateway infrastructure as text and image requests, meaning developers can manage all modalities—voice, text, video—with one API key, unified observability, and shared spend controls. A browser-based playground lets developers test audio models without writing code.
Summaries like this, in your inbox every morning.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
OpenAI began offering a ChatGPT feature that accepts uploaded audio files and can transcribe them, summarize t…

NECネクサソリューションズ started offering "対話型音声AIエージェント Powered by DEN.Ai" on 2026年10月6日

Automation Anywhere announced Wednesday an agreement to acquire Boost.ai, a conversational voice AI company, f…
Willow Care Inc., maker of the AI dictation app Willow Voice, picked CoreWeave Inc
Abu Dhabi's state-run Technology Innovation Institute released Falcon-Emirati models trained on the Emirati Ar…

Google LLC launched the web-based SynthID Detector, which flags AI-made images, video and audio from Google, O…