AIToday
Large Language ModelsOpen-Source AIGoogle DeepMindPublished: Oct 7, 2026, 06:00 JST

EmbeddingGemma 2: 740M-param open model goes multimodal

EmbeddingGemma 2: 740M-param open model goes multimodal

3 Key Points

  1. What happened

    Google DeepMind launched EmbeddingGemma 2, an open multimodal embedding model with 740 million parameters under a commercially permissive Apache 2.0 license, after EmbeddingGemma drew more than 20 million downloads.

  2. Why it matters

    The earlier text-only version already reached more than 20 million downloads, so a version that handles images, audio and video may extend that reach to search tools that until now ran only on text.

  3. What to watch

    Its on-device pitch hinges on the stated memory figures — about 191MB active RAM for text-only weights and about 567MB for the full multimodal model on a Google Pixel 11 Pro. Watch for independent results on the MTEB Code and MAEB benchmarks.

WHO IT HITSDevelopers building on-device and offline apps benefit most, since the model can be downloaded from Hugging Face and Kaggle and deployed through tools such as LiteRT and Google AI Edge MediaPipe. Teams handling sensitive audio, image or video data may find a locally run embedding model attractive, though the memory figures are Google's own.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

EmbeddingGemma arrived last year as a lightweight option for high-quality text embeddings, aimed at helping apps organize, search, and connect information directly on consumer hardware. Google DeepMind says the developer community's response blew past its expectations, with more than 20 million downloads and use in on-device search tools and privacy-first retrieval pipelines. EmbeddingGemma 2 is the follow-up, built on the Gemma 4 architecture and expanding beyond text to unify code, images, video, and audio in a shared embedding space.

The design choices point in one direction: staying small enough to run locally. It has 740 million parameters, but can need as little as 270M parameters for text-only workloads, with optional vision (170M) and audio (300M) encoders. Matryoshka Representation Learning lets developers truncate output vectors from 768 dimensions down to 512, 256, or 128, which Google says provides up to 6x storage reduction for local vector databases. Its 8K token context window, 4x larger than EmbeddingGemma 1, lets it process up to 5.5 minutes of audio, 29 images, 58 video frames, or interleaved combinations.

On quality, Google says EmbeddingGemma 2 leads among sub-1B multimodal embedders on benchmarks like MTEB Code and MAEB, and matches or outperforms many larger models across text, vision, and audio. The outcome hinges on whether that quality-per-parameter claim holds up as developers test it on their own hardware, and on whether the memory figures translate into smooth performance outside Google's own Pixel 11 Pro. Google says Gemini Enterprise Agent Platform Model Garden availability is coming soon.

FAQ
What license does EmbeddingGemma 2 use?
It is released under a commercially permissive Apache 2.0 license.
How much memory does it need on a phone?
With quantization on a Google Pixel 11 Pro, Google says it needs as little as ~191MB active RAM for text-only weights and ~567MB for the full multimodal model.
How much did code performance improve?
It showed a 9.92-point improvement in MTEB Code, from 68.76 to 78.68.
Google DeepMindRead Original Article

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleVast: Alon Horev says AI agent memory moving to tiered storage