Google DeepMind released EmbeddingGemma 2 on October 6, 2026, calling it "our first natively multimodal open model for on-device embeddings." It extends the family beyond text to unify code, images, audio and video in a single shared embedding space. The model has 740M parameters, is released under an Apache 2.0 license, and its weights are on Hugging Face and Kaggle. Details are in Google's launch post.
One space for every modality
An embedding model turns an input into a vector so that similar meanings land close together. In EmbeddingGemma 2, the word "cat," a photo of a cat, an audio clip of one and a video of one are all meant to cluster by meaning, whatever the format.
That makes cross-modal lookups possible without a separate model per format. Google's own example is finding moments in a video using a voice memo. More generally, developers can add multimodal search to an app, or pair the model with Gemma 4 for private, on-device retrieval-augmented generation (RAG). Google's DiffusionGemma is another open model built on a Gemma 4 backbone.
What the three charts show
Google's thread plots mean score against model size on three benchmarks, with a dotted line marking the size-versus-score frontier.
Code retrieval, from Google DeepMind's launch thread and blog post.
Image retrieval, from Google DeepMind's launch thread and blog post.
Audio embeddings, from Google DeepMind's launch thread and blog post.
Reading values off the charts, roughly:
- Code (MTEB Code): EmbeddingGemma 2 scores about 77. That is well above its 300M predecessor, embeddinggemma-300m (about 67), inf-retriever-v1-1.5b (about 67) and granite-embedding-311m (about 63). It is roughly level with Qwen3-Embedding-0.6B (about 75.5) and with pplx-embed-v1-4b (about 77.5), a model more than five times larger. Only Qwen3-Embedding-8B, at about 81, is clearly ahead.
- Image (MIEB Lite): about 65, against about 66 for LCO-Embedding-Omni-3B, about 61.5 for jina-embeddings-v5-omni-small and about 53.5 for SigLIP so400m, which supports only image and text.
- Audio (MAEB): about 48.5, ahead of e5-omni-3B (about 48) and Qwen2-Audio-7B (about 35), but behind jina-embeddings-v5-omni-nano (about 51), BidirLM-Omni-2.5B (about 53) and LCO-Embedding-Omni-7B (about 56).
What the claims support
Google's wording is that the model is "competitive across benchmarks," even outperforming some specialist models more than twice its size. The charts back that up: on all three, EmbeddingGemma 2 sits at or near the frontier line for its size, one 740M model landing close to 3B-class omni models on images and level with a 4B model on code. It is not the top scorer anywhere, though. The leader is Qwen3-Embedding-8B on code, LCO-Embedding-Omni-3B on images and LCO-Embedding-Omni-7B on audio, and audio is the weakest of the three results, with four models scoring higher.
The pitch is breadth and footprint rather than a single record: one small Apache 2.0 model covering five modalities.
Context
Apache 2.0 sits near the permissive end of the ladder in How Open Is 'Open'?, and a 740M model is at the opposite end of the size spectrum from the models in Open Weights You Can't Actually Run. SigLIP, the image baseline in Google's chart, is explained in SigLIP, Explained. For other small models built to run locally, see Desert Ant Labs' on-device models.