Google DeepMind Releases EmbeddingGemma 2, a Local Multimodal Embedding for Text, Image, Audio and Video

GOOGL GOOG

·

Google DeepMind on Tuesday released EmbeddingGemma 2, an Apache-2.0-licensed multimodal embedding model that developers can run locally to generate shared numerical representations for text, images, audio and video frames. The 740 million-parameter model, announced by Google DeepMind and Google AI, is available through the Gemma model card, a Google Developers and Google AI Edge blog post, and on Hugging Face.

The release matters because embedding models are the plumbing behind search, retrieval, classification and clustering systems. They convert content into vectors that software can compare mathematically. A multimodal embedding model goes a step further by placing different types of media into the same vector space, so a text query can retrieve a matching image, audio clip or video moment. Google is explicitly pitching EmbeddingGemma 2 for local, privacy-first applications, including on-device search and retrieval-augmented generation, or RAG, a technique that lets software pull in relevant information during a task.

Google describes the model as “an open multimodal embedding model,” and the model card says, “EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind.” The architecture is split into a 270 million-parameter text stack — made up of a 130 million-parameter backbone and a 140 million-parameter embedder — plus modular encoders for vision at 170 million parameters and audio at 300 million. Developers can load only the parts they need, which Google says brings the footprint to about 270 million parameters for text only, 440 million for text and image, 570 million for text and audio, and 740 million for the full multimodal setup.

All of those inputs are mapped into a single 768-dimensional vector space, including text, code, images, video frames and audio. The model also supports Matryoshka Representation Learning, which lets developers truncate embeddings to 512, 256 or 128 dimensions and then re-normalize them, potentially reducing storage and compute costs in vector search systems. Google says EmbeddingGemma 2 supports an 8,192-token context window and can process minutes of audio or video.

In its launch materials, Google emphasized local deployment. “Designed specifically for local, privacy-first applications, EmbeddingGemma 2 covers text, vision, and audio modalities in a compact 740M parameter footprint,” the company wrote in a Google Developers blog post. Google said that with quantization and selective loading on a Google Pixel 11 Pro, the model can use as little as about 191 MB of active RAM for text-only weights and about 567 MB for the full multimodal model. The company also pointed to demo apps including AI Edge Gallery Instant Media Search, Video Moments Finder and Foresight, a Mac app for local meeting assistance.

The weights and artifacts are published on Hugging Face, where Google included sample usage for developers, including with Sentence Transformers and Transformers. Google also said the model will be available as a service through ML Kit on Android in the coming weeks and is planned for MediaPipe Tasks for cross-platform use, with NPU acceleration where available.

EmbeddingGemma 2 follows the earlier, text-focused EmbeddingGemma released in 2025 and extends the Gemma family into native multimodality and edge deployment. Google also said the model is a pre-trained embedding system with no post-training alignment or output-level moderation. The company said its safety work focused on pre-training data filtering, and that developers should add application-level safeguards and follow the Gemma Prohibited Use Policy. Google published internal benchmark tables in the model card, but those results are company-reported rather than independent validation.

Tags: #ai, #embeddings, #google, #multimodal

Stocks: GOOGL GOOG