AIToday
Large Language ModelsAudio & SpeechSimon Willison's WeblogPublished: Sep 24, 2026, 06:00 JST

Google drops Gemini 3.8 TTS with 2,000 voices

Google drops Gemini 3.8 TTS with 2,000 voices

3 Key Points

  1. What happened

    Google released gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts, offering over 2,000 voices and custom voice creation from a 30-second audio sample.

  2. Why it matters

    Two new voice models give users a wider choice of voices and make it easier to create a custom voice with only a short sample, which could help developers build more natural-sounding applications.

  3. What to watch

    The cost and speed trade-offs between the standard and Flash-Lite versions will determine which one developers pick. In a demo, generating 1m 18s of audio with Gemini 3.8 Flash TTS took about 20 seconds and cost 2.74 cents.

WHO IT HITSDevelopers and product teams building voice-enabled apps now have two cheaper Gemini TTS options with over 2,000 voices and a quick custom-voice path, which may lower the barrier to adding natural speech to their products.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Google's introduction of two new text-to-speech models, gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts, expands its Gemini lineup with options for both standard and lighter-weight speech synthesis. The models ship with a library of over 2,000 voices and allow users to create a custom voice from just a 30-second audio sample, which could make personalized voice generation more accessible.

A key feature is the ability to define multi-speaker conversations with different voices and style instructions, as shown in a demo where two pelicans debate moving to the Pacifica Pier. That demo used Gemini 3.8 Flash TTS (not the cheaper Flash-Lite) and took about 20 seconds to generate 1m 18s of audio at a cost of 2.74 cents, suggesting a potentially practical balance of speed and price for developers.

The real test will be whether the Flash-Lite version delivers acceptable quality at a lower cost, and how developers weigh the trade-offs between the two models for their applications.

FAQ
How many voices are available?
Over 2,000 voices are included.
How much does it cost to generate audio?
A 1m 18s clip using Gemini 3.8 Flash TTS cost 2.74 cents.
Simon Willison's WeblogRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Microsoft Copilot rebuilt with Home, Code, Autopilot tabsITmedia AI+ · 2h ago
  • Salesforce at Dreamforce 2026: UCLA Health hits 75,000+ AI chatsTop Companies AI · 5h ago
  • Kavukcuoglu: Gemini 4 hits post-training, early ship eyedTop Companies AI · 5h ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleOpenAI agents hacked Hugging Face for test answers