AIToday
Large Language ModelsAudio & SpeechAI Business & IndustryITmedia AI+Published: Sep 24, 2026, 19:01 JST

Google's Gemini 3.8 Flash TTS builds voices from scratch, tops Hume AI test

Google's Gemini 3.8 Flash TTS builds voices from scratch, tops Hume AI test

3 Key Points

  1. What happened

    Google announced Gemini 3.8 Flash TTS and Flash-Lite TTS on September 23, which generate new voices from natural-language prompts instead of picking among 30 preset voices, and Flash TTS scored 1st overall (71.4) on Hume AI's voice design benchmark.

  2. Why it matters

    Creators making games, audiobooks, podcasts or voice agents can now specify a character, accent and tone in words rather than choose from a fixed list, which appears to change how quickly a distinct voice can be produced.

  3. What to watch

    The pricing is a limited-time offer — $9 per 1M tokens of audio output for Flash TTS through December 31, 2026, doubling from January 1, 2027 — so budgeting hinges on whether teams lock in usage before then.

WHO IT HITSProducers of games, audiobooks, podcasts and dubbed content, plus developers building voice agents on Gemini API, Google AI Studio, Google Vids and Gemini Notebook, are the ones whose tooling and per-token costs change here.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Google already offered text-to-speech, but its earlier models worked by choosing among 30 preset voices. The new pair changes that: Flash TTS and Flash-Lite TTS let users describe a character, accent or vocal quality in plain language and generate a voice that did not exist before, and any voice created can be saved so it stays consistent across a long project. Google says a "voice remix" feature, which adjusts timbre, pitch, speaking speed and accent of library voices via prompt, is coming soon.

The two models split by job. Flash TTS targets character voices and fine direction for games, audiobooks and podcasts, while the cheaper Flash-Lite TTS targets bulk dubbing, audio content production and voice agents. Language coverage differs too — 130 languages for Flash TTS and 101 for Flash-Lite TTS, both including Japanese, with a library of more than 2000 voices covering regional speech differences.

Google is pairing the capability with controls. Generated audio carries its inaudible SynthID watermark, voice replication adds C2PA provenance credentials and requires a consent recording, and access to replication is restricted in several markets. Flash TTS also placed 1st overall on Hume AI's voice design benchmark (71.4) and 1st on Hume's overall quality metric, with Flash-Lite TTS 2nd. Whether that positioning holds up in real production use is likely to depend on how the consent checks and the doubling prices from January 1, 2027 are absorbed by the teams doing high-volume work.

FAQ
How much does Gemini 3.8 Flash TTS cost?
On the paid Gemini API plan, Flash TTS costs $0.5 per 1M tokens for text input and $9 per 1M tokens for audio output. Flash-Lite TTS is $0.5 for input and $6 for output. Both prices run through December 31, 2026 and double from January 1, 2027.
Where can I use these models, and where is voice replication blocked?
Developers get Flash TTS and Flash-Lite TTS through Gemini API and Google AI Studio; general users get Flash TTS in Gemini Notebook and Flash-Lite TTS in Google Vids. Voice replication in Google AI Studio is not available in Illinois and Texas, the European Economic Area, the UK, Switzerland and India.
Can these models clone my voice?
Yes — a voice replication feature reproduces a voice from a 30-second audio sample, limited to your own voice or one you have the right to use. You must submit a recording of the voice owner giving verbal consent, and the system checks whether the speaker in that recording matches the reference audio.

Also reported by THE DECODER

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Schmidhuber joins Sakana AI to lead RSI LabITmedia AI+ · 46m ago
  • TypeSafe AI launches Jev, claiming up to 200x faster inference without hallucinationsITmedia AI+ · 46m ago
  • Anthropic touts Claude's ART enzyme find; CRISPR researcher says routine genome miningTHE DECODER · 46m ago

AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleAnthropic IPO may slip to November, WSJ says