AIToday
Large Language ModelsAudio & SpeechSimon Willison's WeblogPublished: Sep 16, 2026, 10:01 JST

Google launches Gemini 3.8 Live speech-to-speech models

Google launches Gemini 3.8 Live speech-to-speech models

3 Key Points

  1. What happened

    Google released Gemini 3.8 Live and 3.8 Live Extended Thinking, two speech-to-speech models similar in shape to OpenAI's GPT-Live family.

  2. Why it matters

    Google now offers two tiers of live speech models, one standard and one for extended thinking, matching the shape of OpenAI's GPT-Live lineup.

  3. What to watch

    A developer built a browser demo letting you pick a model and voice preset, enter a system prompt, and interrupt the model mid-speech. Watch whether Google publishes its own tutorial on the WebSocket API.

WHO IT HITSDevelopers building voice interfaces can now choose between two Google speech-to-speech models, one tuned for extended thinking. Tool builders following the WebSocket API tutorial can test them in a browser without extra libraries.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

Google's release of Gemini 3.8 Live and 3.8 Live Extended Thinking adds two speech-to-speech models to its lineup. The pair mirrors the shape of OpenAI's GPT-Live family, with one standard model and one extended-thinking variant, suggesting Google is positioning the two as alternatives for live voice applications rather than as distinct products.

Alongside the release, a browser-based demo UI was built against the documentation, letting users select a model and voice preset, enter an optional system prompt, and hold a voice conversation that can be interrupted while the model is talking. The implementation deliberately avoids libraries, connecting directly to a Google WebSocket endpoint and using a Web Audio API AudioContext for both capture and playback, which points to how lightweight a client for these models can be.

The practical test for these models is likely whether developers can get live voice conversations working without heavy dependencies, and whether Google's own tutorial makes that setup straightforward. For teams building voice interfaces, the extended-thinking variant may matter most where a live answer benefits from more deliberation, though that depends on how the models behave in practice.

FAQ
What models did Google release?
Google released Gemini 3.8 Live and 3.8 Live Extended Thinking, both speech-to-speech models similar in shape to OpenAI's GPT-Live family.
How can I try the new models?
A web UI built for trying the models lets you select a model and voice preset, enter an optional system prompt, and start a voice conversation through your browser, including interrupting the model while it is talking.
What technical approach does the demo use?
The demo uses no libraries. It connects to a Google WebSocket endpoint and uses a Web Audio API AudioContext for both capture and playback.
Simon Willison's WeblogRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Profound raises $180 million at $1.8 billion for AI visibilitySiliconANGLE AI · 1h ago
  • Google's Gemini 3.8 Live brings real-time voice reasoningSiliconANGLE AI · 1h ago
  • At Dreamforce, Salesforce Debuts Koa with Nvidia, Claudeforce with AnthropicSiliconANGLE AI · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleAltman, Amodei AI slowdown plea rejected ahead of Sept. 24