
What happened
Google DeepMind launched EmbeddingGemma 2, an open multimodal embedding model with 740 million parameters under a commercially permissive Apache 2.0 license, after EmbeddingGemma drew more than 20 million downloads.
Why it matters
The earlier text-only version already reached more than 20 million downloads, so a version that handles images, audio and video may extend that reach to search tools that until now ran only on text.
What to watch
Its on-device pitch hinges on the stated memory figures — about 191MB active RAM for text-only weights and about 567MB for the full multimodal model on a Google Pixel 11 Pro. Watch for independent results on the MTEB Code and MAEB benchmarks.
WHO IT HITSDevelopers building on-device and offline apps benefit most, since the model can be downloaded from Hugging Face and Kaggle and deployed through tools such as LiteRT and Google AI Edge MediaPipe. Teams handling sensitive audio, image or video data may find a locally run embedding model attractive, though the memory figures are Google's own.
Summaries like this, in your inbox every morning.
EmbeddingGemma arrived last year as a lightweight option for high-quality text embeddings, aimed at helping apps organize, search, and connect information directly on consumer hardware. Google DeepMind says the developer community's response blew past its expectations, with more than 20 million downloads and use in on-device search tools and privacy-first retrieval pipelines. EmbeddingGemma 2 is the follow-up, built on the Gemma 4 architecture and expanding beyond text to unify code, images, video, and audio in a shared embedding space.
The design choices point in one direction: staying small enough to run locally. It has 740 million parameters, but can need as little as 270M parameters for text-only workloads, with optional vision (170M) and audio (300M) encoders. Matryoshka Representation Learning lets developers truncate output vectors from 768 dimensions down to 512, 256, or 128, which Google says provides up to 6x storage reduction for local vector databases. Its 8K token context window, 4x larger than EmbeddingGemma 1, lets it process up to 5.5 minutes of audio, 29 images, 58 video frames, or interleaved combinations.
On quality, Google says EmbeddingGemma 2 leads among sub-1B multimodal embedders on benchmarks like MTEB Code and MAEB, and matches or outperforms many larger models across text, vision, and audio. The outcome hinges on whether that quality-per-parameter claim holds up as developers test it on their own hardware, and on whether the memory figures translate into smooth performance outside Google's own Pixel 11 Pro. Google says Gemini Enterprise Agent Platform Model Garden availability is coming soon.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Meta, Walmart, Stripe and Sierra Technologies are creating the Personal Agent Protocol, introduced by Sierra c…
Mistral AI opened a public preview of Mistral Large 4, its 1.05 trillion-parameter MoE model nicknamed "Le Cho…

Reflection AI started early access to Beam, its first open-weights model, with 501 billion total and 23 billio…

Google began rolling out "Simple Guide" in Gemini Live on Android, letting users share camera or screen views…

Anthropic launched Claude for Google Workspace as a public beta for paid Claude plans, adding Claude to Google…

Testing a fake LLM call on CPython 3.14.8, the developer saw the first await succeed and the second raise Runt…
