
What happened
Google DeepMind's google/embeddinggemma-2, a roughly 740M-parameter open multimodal embedding model, ran in free Colab on a T4 and embedded text, images and audio into one 768-dimension space.
Why it matters
EmbeddingGemma 2 maps several media types into a single 768-dimension space, so similarity comparison across them becomes possible, though in this small test the results were not equally clean in every direction.
What to watch
The test is a small, dated example, so the result hinges on whether the uneven cross-modal retrieval holds at larger scale; Google DeepMind's model card also recommends bfloat16 or float32 rather than float16.
WHO IT HITSDevelopers and researchers building cross-media search or RAG pipelines can now prototype a single embedding model on a free Colab GPU, but should validate each query direction on their own data rather than assume uniform accuracy.
Summaries like this, in your inbox every morning.
EmbeddingGemma 2 is Google DeepMind's open multimodal embedding model, positioned by the article as taking architectural advances from Gemma 4 rather than being a Gemma 2 variant. It sits in a different role from a large language model: instead of producing answers, it turns data into a 768-dimension vector that can be used for retrieval, similarity scoring, clustering, classification, RAG or recommendation. The article tested it on free Google Colab, where the assigned GPU was a T4, and the 740M-scale model loaded without trouble.
The experiment deliberately used only images with clearly documented licenses, self-written short Japanese texts and audio of known provenance, including LibriSpeech under CC BY 4.0 and Python-generated tones. The results showed that multilingual, cross-media retrieval is not equally strong in all directions. Japanese text queries such as asking for a cat photo or mentioning wanting coffee ranked the right text and image high, but using the cat image itself as a query did not cleanly surface the matching cat description, and a spoken-voice query did not pull up the actual speech clip within the top seven. A separate trial of Matryoshka Representation Learning, which can shorten the embedding to 128, 256, 512 or 768 dimensions, kept the same top-ranked space document at 768 and 256 dimensions in that one example.
The author also raises explainability as an open question: because different encoders all write into one shared representation, it is not obvious why an image and a text land near each other, and the cat-image result is a concrete case. Whether this model becomes practical for a personal cross-media RAG over lecture materials, photos, videos and recordings will likely depend on whether these direction-specific gaps narrow with larger, more varied data. The article notes that its results are as of October 7, 2026, and that Colab's GPU and the relevant libraries may change.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Kore.ai Inc. launched Autoloop, which adjusts AI agents toward goals customers set for task completion, busine…
OpenAI launched a Decisions API that evaluates text, images, or both, returning yes/no probabilities, category…

Google launched Playground, a browser-based AI platform that lets adults in the US create games using only tex…

Google made its SynthID detector available globally to anyone with a Google, OpenAI, or Apple account, with a…

In August, DoorDash emailed Bay Area restaurants warning they may be listed on Bites without consent, with ter…

At MIT Future Fest, Tony Fadell said the Rabbit R1, Humane Ai pin, and Limitless pendant failed because they "…
