
Google released Gemini 3.5 Transcribe, a real-time speech-to-text model.
It supports 85 languages and auto-corrects verbal stumbles.
The model shows better accuracy and 70 percent lower latency than Chirp 3.
What happened
Google has launched Gemini 3.5 Transcribe, a speech-to-text model for real-time transcription that auto-corrects slips of the tongue and strips filler words like "um."
Why it matters
Google reports a word error rate of 4.0 percent for streaming and 2.6 percent for recorded audio, with 70 percent lower latency than its predecessor, Chirp 3. The model can also hand off tasks like image generation or web searches to other Gemini models through "function calling."
What to watch
The model is available in Google AI Studio and on the Gemini Enterprise Agent Platform, and is already built into Gboard for Android and the Gemini app on macOS, with Chrome support coming soon.
Ask the AI about this article →
This launch follows Google's previous speech-to-text work with Chirp 3, and Gemini 3.5 Transcribe directly improves upon it by delivering 70 percent lower latency. The model is positioned for broad integration, already appearing in Gboard for Android and the Gemini app on macOS. Its ability to recognize over 85 languages automatically and format text by itself aims to make transcription more seamless.
The inclusion of "function calling" allows Gemini 3.5 Transcribe to pass tasks like image generation or web searches to other Gemini models. This suggests the transcription tool is designed not just to convert speech to text, but to act as a starting point in a chain of AI-driven actions. With Chrome support coming soon, its availability is set to expand beyond current platforms.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Plaud Inc. introduced the Plaud One Explorer Edition, a pair of earbuds and a charging case that serve as a mo…
Sophus IT Solutions, an AI native engineering and professional services firm, has been named an OpenAI Select…

Z.ai released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series

Relay, founded by two former Nothing employees, is debuting a dedicated microphone for high-fidelity voice-to-…

Nvidia is acquiring open-source AI platform Hugging Face for $12.9 billion, according to The Information

Anthropic is adding a browser directly into Claude Cowork, its desktop app
