AIToday
Large Language ModelsOpen-Source AIZenn AI/MLPublished: Oct 7, 2026, 22:00 JST

google/embeddinggemma-2: one 768-dim space, uneven search

google/embeddinggemma-2: one 768-dim space, uneven search

3 Key Points

  1. What happened

    Google DeepMind's google/embeddinggemma-2, a roughly 740M-parameter open multimodal embedding model, ran in free Colab on a T4 and embedded text, images and audio into one 768-dimension space.

  2. Why it matters

    EmbeddingGemma 2 maps several media types into a single 768-dimension space, so similarity comparison across them becomes possible, though in this small test the results were not equally clean in every direction.

  3. What to watch

    The test is a small, dated example, so the result hinges on whether the uneven cross-modal retrieval holds at larger scale; Google DeepMind's model card also recommends bfloat16 or float32 rather than float16.

WHO IT HITSDevelopers and researchers building cross-media search or RAG pipelines can now prototype a single embedding model on a free Colab GPU, but should validate each query direction on their own data rather than assume uniform accuracy.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

EmbeddingGemma 2 is Google DeepMind's open multimodal embedding model, positioned by the article as taking architectural advances from Gemma 4 rather than being a Gemma 2 variant. It sits in a different role from a large language model: instead of producing answers, it turns data into a 768-dimension vector that can be used for retrieval, similarity scoring, clustering, classification, RAG or recommendation. The article tested it on free Google Colab, where the assigned GPU was a T4, and the 740M-scale model loaded without trouble.

The experiment deliberately used only images with clearly documented licenses, self-written short Japanese texts and audio of known provenance, including LibriSpeech under CC BY 4.0 and Python-generated tones. The results showed that multilingual, cross-media retrieval is not equally strong in all directions. Japanese text queries such as asking for a cat photo or mentioning wanting coffee ranked the right text and image high, but using the cat image itself as a query did not cleanly surface the matching cat description, and a spoken-voice query did not pull up the actual speech clip within the top seven. A separate trial of Matryoshka Representation Learning, which can shorten the embedding to 128, 256, 512 or 768 dimensions, kept the same top-ranked space document at 768 and 256 dimensions in that one example.

The author also raises explainability as an open question: because different encoders all write into one shared representation, it is not obvious why an image and a text land near each other, and the cat-image result is a concrete case. Whether this model becomes practical for a personal cross-media RAG over lecture materials, photos, videos and recordings will likely depend on whether these direction-specific gaps narrow with larger, more varied data. The article notes that its results are as of October 7, 2026, and that Colab's GPU and the relevant libraries may change.

FAQ
What is EmbeddingGemma 2 used for?
It converts text, images, audio and video into a 768-dimension vector for similarity comparison and retrieval, not for generating answers. The article describes it as useful for search, clustering, classification, RAG and recommendations.
What are its main specs?
The article lists about 740M total parameters: 270M for text, 170M for vision and 300M for audio. It outputs 768 dimensions, supports 100+ languages, has an 8,192-token context window and is Apache 2.0 licensed.
Which direction of cross-media search worked least well?
In the article's small test, Image→Text performed poorly: the cat image put the coffee image and camera image ahead of the matching cat text. Text→Audio also failed to place the LibriSpeech speech clip in the top seven.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleChatGPT forges signatures of over 15 New Yorker cartoonists