Google has launched EmbeddingGemma 2, an open embedding model that can map text, code, images, video and audio into a single shared space, and is small enough to run on a phone. The 740-million-parameter model is released under the Apache 2.0 license, and is meant for developers building search and retrieval tools that work entirely on-device.

The model is built on the architecture of Gemma 4, the open model family Google released in April. Google says it also draws on the same technology as its Gemini Embedding models.
What is new
Embedding models turn content into numerical vectors so that software can search and compare it by meaning rather than by keywords. The first EmbeddingGemma, released last year, handled text only. Google says it has been downloaded more than 20 million times, with developers using it for on-device search and privacy-focused retrieval augmented generation (RAG) pipelines.
EmbeddingGemma 2 extends that to multiple modalities. Google says a single model can, for instance, find a specific video clip from a voice memo, or search through hours of audio recordings using a text query.
The model is modular. Text-only workloads need as little as 270 million parameters, with optional vision (170 million) and audio (300 million) encoders that can be added for full multimodal support. The context window has grown to 8,000 tokens, four times that of the original, enough for up to 5.5 minutes of audio, 29 images or 58 video frames, or combinations of these.
Using Matryoshka Representation Learning, developers can shrink the output vectors from 768 dimensions to 512, 256 or 128, which Google says cuts storage for local vector databases by up to 6x.
Performance
Google says EmbeddingGemma 2 leads sub-1B multimodal embedders on benchmarks including MTEB Code and the Massive Audio Embedding Benchmark (MAEB), and matches or beats many larger models across text, vision and audio tasks. It adds that the model even outperforms some specialist models more than twice its size.
Text performance is on par with the original model, but code retrieval sees a big jump: the MTEB Code score rises from 68.76 to 78.68. Google says this makes the model well suited to local codebase indexing, semantic code search and retrieval for coding agents.
Built to run on a phone
With quantization, Google says the model needs about 191MB of active RAM for text-only weights and around 567MB for the full multimodal version on a Pixel 11 Pro. Because it shares a text tokenizer and audio encoder with Gemma 4, developers can run the two models together in one pipeline with a lower combined memory footprint, for example pairing EmbeddingGemma 2 for local file retrieval with Gemma 4 for reasoning over what it finds.
On-device embeddings are also attractive for privacy, since user data never has to leave the device. That fits Google’s broader push around on-device AI, which gained visibility when the Google AI Edge Gallery app climbed into the App Store’s top 10 on the back of Gemma 4. Google is showing off EmbeddingGemma 2 in that app through features like Instant Media Search and Video Moments Finder, and in its Google AI Edge Foresight app for local file retrieval.
Availability
The model weights are on Hugging Face and Kaggle, with availability on Google’s Model Garden coming soon. Google says it has worked with partners so that the model runs with tools including transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama and LMStudio, and in the browser through transformers.js and WebGPU. Qdrant is supported for storing vectors, and Unsloth has published guidance on fine-tuning the model.
For on-device deployment, developers can use Google AI Edge’s MediaPipe and LiteRT.
The launch gives developers an open, commercially usable option for multimodal retrieval that can run locally, rather than relying on a cloud embedding API. Google’s own Gemini Embedding 2, which handles similar modalities, remains available as a hosted model.