
What happened
Google announced Gemini 3.8 Flash TTS and Flash-Lite TTS on September 23, which generate new voices from natural-language prompts instead of picking among 30 preset voices, and Flash TTS scored 1st overall (71.4) on Hume AI's voice design benchmark.
Why it matters
Creators making games, audiobooks, podcasts or voice agents can now specify a character, accent and tone in words rather than choose from a fixed list, which appears to change how quickly a distinct voice can be produced.
What to watch
The pricing is a limited-time offer — $9 per 1M tokens of audio output for Flash TTS through December 31, 2026, doubling from January 1, 2027 — so budgeting hinges on whether teams lock in usage before then.
WHO IT HITSProducers of games, audiobooks, podcasts and dubbed content, plus developers building voice agents on Gemini API, Google AI Studio, Google Vids and Gemini Notebook, are the ones whose tooling and per-token costs change here.
Summaries like this, in your inbox every morning.
Google already offered text-to-speech, but its earlier models worked by choosing among 30 preset voices. The new pair changes that: Flash TTS and Flash-Lite TTS let users describe a character, accent or vocal quality in plain language and generate a voice that did not exist before, and any voice created can be saved so it stays consistent across a long project. Google says a "voice remix" feature, which adjusts timbre, pitch, speaking speed and accent of library voices via prompt, is coming soon.
The two models split by job. Flash TTS targets character voices and fine direction for games, audiobooks and podcasts, while the cheaper Flash-Lite TTS targets bulk dubbing, audio content production and voice agents. Language coverage differs too — 130 languages for Flash TTS and 101 for Flash-Lite TTS, both including Japanese, with a library of more than 2000 voices covering regional speech differences.
Google is pairing the capability with controls. Generated audio carries its inaudible SynthID watermark, voice replication adds C2PA provenance credentials and requires a consent recording, and access to replication is restricted in several markets. Flash TTS also placed 1st overall on Hume AI's voice design benchmark (71.4) and 1st on Hume's overall quality metric, with Flash-Lite TTS 2nd. Whether that positioning holds up in real production use is likely to depend on how the consent checks and the doubling prices from January 1, 2027 are absorbed by the teams doing high-volume work.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Sakana AI said Jürgen Schmidhuber is joining as chief scientific advisor and will lead its new RSI Lab, which…

On 2026年9月15日, TypeSafe AI published System One Model and opened early access to Jev, which returns only type-…

Microsoft's Azure revenue grew 43%, Microsoft Cloud revenue rose 27% to $59.3 billion, and Amazon's AWS second…

CrowdStrike trades near 177 times forward earnings and 167 times free cash flow, while Palo Alto Networks trad…

Nvidia CEO Jensen Huang called fears over AI's existential threats a "distraction" on The Ezra Klein Show, as…

Anthropic said Claude largely on its own found an unknown enzyme system, "ART," in DNA databases
