AIToday
Large Language ModelsAudio & SpeechArs Technica AIPublished: Aug 27, 2026, 06:01 JST2 min read

Google unveils Gemini 3.5 Transcribe for cleaner voice input

Google unveils Gemini 3.5 Transcribe for cleaner voice input

Key takeaway

  • Google has announced Gemini 3.5 Transcribe, a new AI speech-to-text model.

  • It is faster and more accurate than previous Chirp 3.

  • The model cleans up verbal stumbles and works in 85 languages.

3 Key Points

  1. What happened

    Google announced Gemini 3.5 Transcribe, an AI model that edits out “ums” and corrections to output polished text. It already powers Gboard’s “Rambler” feature on the Pixel 11 and will appear across the Google ecosystem.

  2. Why it matters

    The new model is about 70 percent faster from voice to final text and has a 5.5 percent live-speech error rate, compared with Chirp 3’s 7.32 percent. It also works in 85 languages and with up to three speakers in pre-recorded audio.

  3. What to watch

    Gemini 3.5 Transcribe can edit text on the fly, refer to your custom vocabulary for specialized jargon, and it may change your wording. This could be unsuitable for situations requiring exact transcription.

Ask the AI about this article →

Context & Analysis

Google’s new Gemini 3.5 Transcribe marks a step forward in voice input technology, promising significant speed gains and modest accuracy improvements over its predecessor, Chirp 3. The model’s ability to clean up speech by removing verbal stumbles and correcting on the fly positions it as a tool for producing polished text directly from voice, rather than just transcribing exactly what was said. This capability, while convenient for casual use, raises questions about fidelity in contexts where exact wording matters, as the AI technically changes the user’s original phrasing.

The model’s integration into Gboard’s Rambler feature on Pixel 11 and its upcoming availability across Google’s ecosystem suggest a broad rollout. With support for 85 languages and up to three speakers in pre-recorded audio, it aims to handle a wide range of voice input scenarios. As with any AI that interprets intent, users will need to weigh the benefits of faster, cleaner output against the potential for altered transcripts in professional or legal settings.

FAQ

When is Gemini 3.5 Transcribe available?
It already powers the Gboard “Rambler” feature on the Pixel 11 and is about to appear throughout the Google ecosystem.
What does Gemini 3.5 Transcribe do differently?
It removes “ums” and “uhs,” edits text on the fly, and uses custom vocabulary for specialized jargon. It works with up to three speakers in pre-recorded audio.
How accurate is Gemini 3.5 Transcribe compared with Chirp 3?
Gemini 3.5 Transcribe has a 5.5 percent live-speech error rate, slightly better than Chirp 3’s 7.32 percent.
Ars Technica AIRead Original Article

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleOpenAI agents hacked Hugging Face due to training flaws