Skip to content
important New models

Google released EmbeddingGemma 2 for local search across text, images and audio

clearly official source

Google released the open model EmbeddingGemma 2, which converts text, code, images, video and audio into a shared space of numerical vectors. It has 740 million parameters, uses the Apache 2.0 license and can run locally without an API key.

Google released EmbeddingGemma 2, a model for searching and comparing content across text, code, images, video and audio. The model converts content into numerical vectors, known as embeddings, in a shared space. The weights are available on Hugging Face and Kaggle under the Apache 2.0 license, which permits commercial use. The model can run locally without an API key; combined with generative models, such as those in the Gemma 4 family, it can retrieve supporting material for answers in applications that work offline.

The full multimodal variant has 740 million parameters. Text requires only the component with 270 million parameters, while the image component adds 170 million and the audio component 300 million. According to Google, the model has a context window of 8 thousand tokens, four times that of EmbeddingGemma 1. It can process up to 5.5 minutes of audio, 29 images, 58 video frames or combinations of these inputs.

According to Google, the output vectors can be shortened from 768 to 512, 256 or 128 dimensions, which can reduce vector storage and memory requirements by up to a factor of six. The company also reports an improvement in the MTEB Code score from 68.76 to 78.68, an increase of 9.92 points, and results comparable to or better than those of some larger models. These results were presented by the manufacturer.

The memory figures in the sources differ in scope: the article in The Decoder cites approximately 191 MB of RAM in general, while Google applies this figure only to active memory for quantized text weights on a Google Pixel 11 Pro device. For the full multimodal model under the same conditions, Google reports approximately 567 MB. The article in The Decoder also reports a query processing time of approximately 20 to 70 milliseconds when using WebGPU in the browser.

What changed

Why it matters

A shared space for different types of content makes it possible, for example, to search for audio recordings with a text query or for a video clip using a voice note. Running the model locally allows developers to process this data without sending it to external servers; the choice of a text-only or multimodal variant affects memory requirements.

Release card

EmbeddingGemma 2

Google

open weights
Specifications
The full multimodal variant has 740 million parameters.
Context
The context window is 8 thousand tokens.
Inputs
The model converts text, code, images, video and audio into numerical vectors in a shared space.
Licence
Apache 2.0
Availability
The weights are available to download on Hugging Face and Kaggle. The model can run locally without an API key.
Documented measurements
  • MTEB Code 78,68 bodu Google reports a score of 78.68 points in a measure of code representation quality, compared with 68.76 points for the previous model, EmbeddingGemma.
According to the sources, it is suitable for
  • Enables local code indexing and semantic search.
  • Enables video search using a voice note or audio recording search using a text query.
  • Combined with a generative model, it enables retrieval of supporting material for answers in applications that work offline.

According to Google, the model retains the multilingual text performance of EmbeddingGemma and improves results for code. The company also reports results comparable to or better than those of some larger models on tasks involving text, images and audio.

The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. It does not yet have a dedicated editorial profile. Model selection and other announcements →

What was added since the original report

Verified updates

  1. New verified information

    Google reports memory usage of approximately 191 MB of RAM.; According to the article, a query using WebGPU in the browser takes approximately 20–70 ms.; According to the article, the model reduces the storage space needed for a local vector database by up to a factor of six.; Google claims that the model outperforms competing models up to twice its size in multimodal benchmarks.

    • Google reports memory usage of approximately 191 MB of RAM.
    • According to the article, a query using WebGPU in the browser takes approximately 20–70 ms.
    • According to the article, the model reduces the storage space needed for a local vector database by up to a factor of six.
    • Google claims that the model outperforms competing models up to twice its size in multimodal benchmarks.

Two audiences, two different impacts

What this means

01

For individuals

A developer can use the model to index their own code locally and search by meaning, without an API key for an external service.

What to do Try indexing your own code locally using the text-only variant of the model and weights from Hugging Face or Kaggle.
More practical updates →
02

For a business

A company can build internal search across documents, images and recordings with local data processing. The Apache 2.0 license permits commercial use, but deployment planning needs to distinguish between the memory requirements of the text-only and full multimodal variants.

Development
What to decide Before integration, check the memory requirements of the variant needed for your data types on the target device.
More business impacts →
Apache 2.0 EmbeddingGemma 2 Gemma 4 Google Google DeepMind Hugging Face MTEB Code RAG WebGPU

Check the original

Event sources

clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.

2
The Decoder (daily AI news) independent context · first detected Google claims EmbeddingGemma 2 outperforms rival embedding models twice its size Google DeepMind Blog primary source EmbeddingGemma 2: an open, lightweight multimodal embedding model