
What happened
Google announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15 and began offering them the same day, rolling them out in Search Live, Gemini Live, Google Workspace, the Gemini API and Gemini Enterprise.
Why it matters
The two models let voice assistants keep talking while tasks run in the background, which could make AI voice agents more usable for hands-on or customer-facing work.
What to watch
Extended Thinking's #1 ranking on Artificial Analysis's Speech to Speech Quality Index uses that index's own scoring; Workspace access requires a paid plan, and API usage pricing is not stated in the official blog.
WHO IT HITSEnterprises evaluating voice agents for customer service and Google Workspace administrators buying Gemini plans will need to check which edition their users get, since Workspace access requires a paid plan and API pricing is not stated.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
Google's last Live-family voice model was Gemini 3.1 Flash Live, released in March, and the version number now matches Gemini 3.8 Flash from September 2. The main change is that 3.8 Live processes video input almost in real time and automatically detects and switches between 97 languages, including Japanese, mid-conversation. Tool execution and API calls run behind the conversation, so the model can acknowledge a request out loud and keep talking before the work finishes; Extended Thinking goes further by speaking and reasoning at once, filling gaps with phrases like "Let me check" and describing multi-step progress by voice.
The rollout spans consumer and business surfaces at once. Regular users get 3.8 Live in Search Live and Extended Thinking in Gemini Live, while Google Workspace brings Extended Thinking to Docs for Google AI Pro and Ultra subscribers and to Gmail and Keep for all Google AI plans. Developers can use the models through the Gemini API and Google AI Studio, and partners such as LiveKit, Pipecat, Agora, Vercel and LangChain can offer the Live API through their platforms. Every generated audio clip carries the SynthID watermark.
The bet here is that voice becomes a practical interface for longer tasks, not just chat. That hinges on whether the latency gains hold up in real deployments, and on how enterprises accept the undecided piece: API usage pricing is not spelled out in the official blog, so cost planning for high-volume voice agents remains an open question.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
At its Amplify user conference, Workiva placed controls, data lineage and human sign-off at the center of its…
At HP Partner Communication 2026 on July 16, 2026, Japan HP's Hiroshi Katsuya told partners to stop selling cl…

The heads of leading U.S

Nikkei's piece argues AI should serve education as a dialogue partner for "wall-bashing" (intellectual sparrin…

At Dreamforce in San Francisco, Anthropic's Dario Amodei reiterated his call for AI developers to slow down, w…
Profound raised $180 million at a $1.8 billion valuation, in a Series D led jointly by Sequoia Capital and Kle…