AIToday
Large Language ModelsAudio & SpeechSiliconANGLE AIPublished: Sep 24, 2026, 10:01 JST

Google's Gemini 3.8 Flash TTS tops Hume AI benchmark

Google's Gemini 3.8 Flash TTS tops Hume AI benchmark

3 Key Points

  1. What happened

    Google launched Gemini 3.8 Flash TTS and Flash-Lite TTS on its cloud platform, with Flash TTS supporting 130 languages and Flash-Lite TTS 101.

  2. Why it matters

    Developers can build audiobooks and other voice products with over 2,000 prepackaged voices and custom voice creation from a 30-second sample, likely lowering barriers to entry.

  3. What to watch

    Flash TTS costs more for better audio quality, so the test is whether buyers accept that trade-off. Watch whether Google adds a third customization option for modifying prepackaged voices.

WHO IT HITSApp developers and content creators building voice features, such as audiobook or video narration, get two new options with different language and price points. Broad deployment may hinge on how buyers weigh Flash TTS's higher price against its better quality.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Google's new text-to-speech models arrive as part of a broader lineup of audio processing tools the company has already released, including algorithms optimized for voice agents, transcription and translation. That history suggests the company is building out a suite rather than a single product, and the two new models share a highly similar application programming interface so developers can use them side-by-side.

A notable feature is customization. Beyond a library of more than 2,000 prepackaged voices, developers can create custom voices through natural language prompts or a 30-second audio sample, with Google requiring consent from the speaker before generating a replica. Google also plans a third option to modify prepackaged voices, extending how much control users have over an AI speaker's timbre, accent and pacing.

The models' strong showing on Hume AI's benchmark may help Google attract developers weighing voice quality against cost, though Flash TTS's higher price and narrower language coverage than Flash-Lite's 101 may influence which model enterprises pick for specific use cases such as audiobooks.

FAQ
How do the two new Google TTS models differ?
Flash TTS supports 130 languages and offers better audio quality for a higher price, while Flash-Lite TTS supports 101 languages and is optimized for cost efficiency and inference speed.
How can developers create custom voices with these models?
They can use natural language prompts, or generate a voice from a 30-second audio sample after securing the speaker's consent. Google also plans a third option for modifying prepackaged voices.
How does Google prevent misuse of AI-generated speech?
Google embeds a SynthID audio watermark that is inaudible to humans but detectable by AI tools, and attaches a C2PA record to every generated audio file.
SiliconANGLE AIRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Microsoft Copilot rebuilt with Home, Code, Autopilot tabsITmedia AI+ · 2h ago
  • Salesforce at Dreamforce 2026: UCLA Health hits 75,000+ AI chatsTop Companies AI · 5h ago
  • Kavukcuoglu: Gemini 4 hits post-training, early ship eyedTop Companies AI · 5h ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleGMI Technology says AI compute leasing sells out in US