
What happened
Google is rolling out Gemini 3.8 Flash TTS and Flash-Lite TTS through the Gemini API and Google AI Studio, supporting more than 100 languages, a 2,000-plus preset voice library, and a cloning feature built from a 30-second audio sample.
Why it matters
Flash TTS is aimed at creative projects such as game characters, audiobooks and podcasts, while Flash-Lite TTS targets low-cost speech generation at scale for dubbing, audio content and voice agents, according to Google. That suggests Google is courting both creators and high-volume business users with the same model family.
What to watch
Google hasn't listed regional endpoints for the new models yet, and Voice Remixing, which would let users adjust timbre, pitch, tempo and accent of library voices, isn't available yet.
WHO IT HITSDevelopers and product teams building voice agents, dubbing pipelines, audiobooks or game characters can now design voices from a text prompt or a 30-second sample through the Gemini API, though EU data-processing endpoints have not been listed for these models.
Summaries like this, in your inbox every morning.
Google's push into speech generation splits its new models by audience. Gemini 3.8 Flash TTS is aimed at creative projects such as game characters, audiobooks and podcasts, while Gemini 3.8 Flash-Lite TTS targets low-cost speech generation at scale for dubbing, audio content and voice agents. Both models are being rolled out through the Gemini API and Google AI Studio. Flash TTS also appears in Gemini Notebook, while Flash-Lite TTS is available in Google Vids, with access through the Gemini Enterprise API to follow.
On the safety side, Google requires anyone whose voice is being cloned to record a spoken statement of consent, and the voice in that recording must match the 30-second sample. Every generated clip also carries an inaudible SynthID watermark. Developer platforms including Agora, LiveKit, Pipecat and Vercel already support integration through the Gemini API, though Google hasn't listed regional endpoints for the new models yet, even as earlier TTS models offered EU data processing.
The test will be whether Google can deliver its promised audio quality at scale. A hands-on test of a preset voice with a style prompt produced a convincing German accent but also a high-pitched background whine and, in one clip, a voice change at the end, which suggests the models may still need refinement before they are ready for demanding production work. Google also says data from the free tier is used to improve its products, while data from the paid tier isn't, which may matter to businesses weighing cost against data handling. And the scheduled price increase on January 1, 2027, from $0.81 to $1.62 per hour of output with Flash TTS, could factor into long-term cost planning for high-volume users.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
ChatGPT Voice now runs on OpenAI's new GPT-6 Astra, Sol, and Luna models and can access plugins like email, ca…

At its Made On YouTube event, YouTube announced Ask Music, a conversational tool built into the YouTube Music…

NVIDIA released Nemotron 3 Diarization, an open-weight, 100M-parameter model that ranks #1 on VoiceArena's Dia…

Alibaba's Qwen released Qwen-Audio-3.1, five models for speech recognition, text-to-speech and real-time inter…

A weekly roundup covers five new generative AI technologies, including YuE2, a music generator compared with S…

Tencent's Hunyuan Speech team and university researchers introduced Gander, which splits real-time conversatio…
