Google released EmbeddingGemma 2 for local search across text, images and audio
Google released the open model EmbeddingGemma 2, which converts text, code, images, video and audio into a shared space of numerical vectors. It has 740 million parameters, uses the Apache 2.0 license and can run locally without an API key.
Google released EmbeddingGemma 2, a model for searching and comparing content across text, code, images, video and audio. The model converts content into numerical vectors, known as embeddings, in a shared space. The weights are available on Hugging Face and Kaggle under the Apache 2.0 license, which permits commercial use. The model can run locally without an API key; combined with generative models, such as those in the Gemma 4 family, it can retrieve supporting material for answers in applications that work offline.
The full multimodal variant has 740 million parameters. Text requires only the component with 270 million parameters, while the image component adds 170 million and the audio component 300 million. According to Google, the model has a context window of 8 thousand tokens, four times that of EmbeddingGemma 1. It can process up to 5.5 minutes of audio, 29 images, 58 video frames or combinations of these inputs.
According to Google, the output vectors can be shortened from 768 to 512, 256 or 128 dimensions, which can reduce vector storage and memory requirements by up to a factor of six. The company also reports an improvement in the MTEB Code score from 68.76 to 78.68, an increase of 9.92 points, and results comparable to or better than those of some larger models. These results were presented by the manufacturer.
The memory figures in the sources differ in scope: the article in The Decoder cites approximately 191 MB of RAM in general, while Google applies this figure only to active memory for quantized text weights on a Google Pixel 11 Pro device. For the full multimodal model under the same conditions, Google reports approximately 567 MB. The article in The Decoder also reports a query processing time of approximately 20 to 70 milliseconds when using WebGPU in the browser.
Why it matters
A shared space for different types of content makes it possible, for example, to search for audio recordings with a text query or for a video clip using a voice note. Running the model locally allows developers to process this data without sending it to external servers; the choice of a text-only or multimodal variant affects memory requirements.
Release card
EmbeddingGemma 2
- Specifications
- The full multimodal variant has 740 million parameters.
- Context
- The context window is 8 thousand tokens.
- Inputs
- The model converts text, code, images, video and audio into numerical vectors in a shared space.
- Licence
- Apache 2.0
- Availability
- The weights are available to download on Hugging Face and Kaggle. The model can run locally without an API key.
- MTEB Code 78,68 bodu Google reports a score of 78.68 points in a measure of code representation quality, compared with 68.76 points for the previous model, EmbeddingGemma.
- Enables local code indexing and semantic search.
- Enables video search using a voice note or audio recording search using a text query.
- Combined with a generative model, it enables retrieval of supporting material for answers in applications that work offline.
According to Google, the model retains the multilingual text performance of EmbeddingGemma and improves results for code. The company also reports results comparable to or better than those of some larger models on tasks involving text, images and audio.
The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. It does not yet have a dedicated editorial profile. Model selection and other announcements →
What was added since the original report
Verified updates
-
Google reports memory usage of approximately 191 MB of RAM.; According to the article, a query using WebGPU in the browser takes approximately 20–70 ms.; According to the article, the model reduces the storage space needed for a local vector database by up to a factor of six.; Google claims that the model outperforms competing models up to twice its size in multimodal benchmarks.
- Google reports memory usage of approximately 191 MB of RAM.
- According to the article, a query using WebGPU in the browser takes approximately 20–70 ms.
- According to the article, the model reduces the storage space needed for a local vector database by up to a factor of six.
- Google claims that the model outperforms competing models up to twice its size in multimodal benchmarks.
Two audiences, two different impacts
What this means
For individuals
A developer can use the model to index their own code locally and search by meaning, without an API key for an external service.
For a business
A company can build internal search across documents, images and recordings with local data processing. The Apache 2.0 license permits commercial use, but deployment planning needs to distinguish between the memory requirements of the text-only and full multimodal variants.
DevelopmentCheck the original
Event sources
clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.