Vercel's AI Gateway now supports real-time voice conversations, text-to-speech, and transcription alongside its existing text and image capabilities.
Real-time voice agents can hold natural conversations where users interrupt and talk over the model, unlike traditional pipelines that chain separate speech-to-text, language model, and text-to-speech steps.
All audio calls route through the same unified API gateway with shared observability and spend controls.
What happened
Vercel's AI Gateway now supports audio and voice capabilities including real-time voice conversations, text-to-speech, and speech-to-text transcription. These features launch with models from OpenAI and xAI, available in beta through AI SDK 7.
Why it matters
Real-time voice agents let applications hold live conversations where users can interrupt and talk over the model as they would with a person, making it practical for voice assistants, customer support, and hands-free tools. Unlike chaining separate speech-to-text, language model, and text-to-speech steps, a single real-time model hears and produces audio directly.
What to watch
Audio calls route through the same AI Gateway infrastructure as text and image requests, meaning developers can manage all modalities—voice, text, video—with one API key, unified observability, and shared spend controls. A browser-based playground lets developers test audio models without writing code.
Ask the AI about this article →
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Apple Music will add “Made With AI” labels to songs where “a material portion” was created using AI, starting…

A software engineer created Schmaudio, a platform that generates interactive audio stories where listeners mak…

Researchers tested 11 widely used open-source speech recognition models and found that several of the highest-…

Apple researchers applied an iterative pseudo-labeling training approach to Mandarin-English code-switching AS…

Adobe is releasing three AI audio tools — Generate Music (royalty-free music for videos), Generate Speech (scr…

Adobe announced general availability of audio generation capabilities in Firefly, its creative AI suite