
What happened
Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its most advanced voice models yet, with Extended Thinking scoring 82.6 on the Artificial Analysis Speech to Speech Quality Index.
Why it matters
Google says the models cut the latency problem in voice AI by running third-party tool calls in the background mid-conversation, so agents can act while still talking, rather than pausing the chat.
What to watch
Extended Thinking's benchmark lead now hinges on whether enterprise users adopt it; watch the $0.005 per minute audio input price and $0.018 per minute output price as the models roll out via the Gemini API.
WHO IT HITSProduct and engineering teams building voice assistants and customer-service agents gain an option that can call tools mid-call instead of stopping the conversation, while enterprise IT teams evaluating Google Workspace and Gemini Enterprise previews will weigh the per-minute audio pricing.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
Google is framing its new voice models around a specific complaint about voice AI agents: the awkward pause when the assistant stops talking to go look something up. Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are designed to push tool and API calls into the background while the conversation keeps its natural pace, so an agent can acknowledge a request with an early verbal cue like "let me check that" and keep the exchange going.
The launch leans heavily on benchmark positioning. Google says Gemini 3.8 Live Extended Thinking hit 82.6 on the Artificial Analysis Speech to Speech Quality Index, surpassing GPT-Live-1-Astra and Grok Voice Think Fast 2.0, while Gemini 3.8 Live placed second on the Speech Agent Arena benchmark and first on ServiceNow's EVA-Bench. The models also carry practical touches for developers: support for 97 languages, near real-time visual grounding, integration through partner platforms such as Vercel, Agora, LiveKit, Pipecat, Fishjam and Vision Agents, and a SynthID watermark on generated audio files to help detect misinformation.
What happens next likely hinges on adoption rather than benchmarks. The models reach developers through the Gemini API and Google AI Studio, with an enterprise private preview in Gemini Enterprise and Search Live, and pricing set at $0.005 per minute for audio inputs and $0.018 per minute for outputs. If those numbers and the background tool-calling hold up in real deployments, Google's voice stack may become a more attractive default for teams building conversational agents, though the body offers no evidence yet on how customers respond.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Keri Tracy, VP and chief audit executive at Newell Brands, told theCUBE that audit has a seat at the table for…
Profound raised $180 million in a Series D led jointly by Sequoia Capital and Kleiner Perkins, at a $1.8 billi…
AIUC raised a $40 million Series A led by Ribbit Capital to start auditing frontier AI models, after previousl…
Factory announced it raised $200 million, backed by Blackstone, Khosla Ventures, Sequoia Capital, NEA and othe…
Salesforce launched Koa, a reasoning model built on Nvidia's Nemotron open-weights platform, and Claudeforce…
ITRI held its ITRI Innovation Day on September 15 in Tainan, Taiwan, gathering research groups and startup tea…
