Google Launches EmbeddingGemma 2: Open Multimodal On-Device Model

Google has launched EmbeddingGemma 2, an open, natively multimodal embedding model designed to run locally on devices. With 740 million parameters, the model maps text, code, images, audio, and video into a single 768-dimensional vector space. It is released under an Apache 2.0 license, enabling offline, privacy-first AI retrieval without cloud dependency.

The 740 Million Parameter Architecture

Google has pushed its on-device AI strategy further by releasing EmbeddingGemma 2. The model is built on the architecture of Gemma 4, which the company previously deployed for local tasks on hardware ranging from smartphones to Raspberry Pi boards. By mapping diverse data types—text, code, images, video, and audio—into a unified 768-dimensional space, the model allows software to perform semantic searches without ever transmitting data to a server.

The efficiency of the model stems from its modular design. Users do not need to load the entire 740-million-parameter weight set if they only require specific modalities. According to Google’s CEO Sundar Pichai, Google’s CEO Sundar Pichai highlighted that the model is designed for on-device efficiency, allowing for a “modular memory footprint.”

Google launches EmbeddingGemma 2, an open multimodal embedding model for de (2026)
  • Text and Code: Requires 270 million parameters.
  • Text, Code, and Vision: Requires 440 million parameters.
  • Text, Code, and Audio: Requires 570 million parameters.
  • Full Multimodal (Text, Code, Vision, Audio, Video): Requires 740 million parameters.

Reducing the Vector Storage Burden

EmbeddingGemma 2 addresses this through Matryoshka Representation Learning (MRL), a technique that allows developers to truncate the 768-dimensional vectors down to 128 dimensions. This reduction can lower storage requirements by up to sixfold while maintaining high retrieval accuracy.

In practice, this means a Pixel 11 Pro can handle text-based embedding tasks using approximately 191MB of memory. When all five modalities are enabled, the memory footprint increases to roughly 567MB. Because the model shares a tokenizer and audio encoder with the existing Gemma 4, developers building retrieval-augmented generation (RAG) pipelines can see significant memory savings by running both models in tandem.

Google Launches EmbeddingGemma 2: Open Multimodal On-Device Model
Photo: timesofindia.indiatimes.com

Privacy Implications of Local Retrieval

By keeping the embedding process entirely on the device, Google is positioning this release as a solution for sensitive data handling. This approach addresses European data protection concerns by eliminating the need to transfer personal audio or visual data to a cloud provider. The model effectively creates a “retrieval pipeline with no network in it,” allowing for the processing of sensitive conversations or private media libraries without external exposure.

However, the lack of safety tuning is a notable trade-off. Google has confirmed that the model has undergone no specific safety tuning or output moderation. Instead, the company states that mitigations have been applied directly to the training data. While the model supports over 100 languages, Google notes that performance parity across these languages is not guaranteed.

EmbeddingGemma 2: what's new in Google's on-device embedding model?

Deployment in the Competitive Ecosystem

The Gemma family has generated over a billion downloads, and the first-generation EmbeddingGemma has had 20 million downloads, enabling the company to build upon its existing developer base to set the standard for on-device retrieval.

Sundar Pichai stated that the model outperforms some specialist models more than twice its size.

The model’s 8,192-token context window is a fourfold increase over its predecessor, providing enough room to process 29 images, 58 video frames, or approximately five and a half minutes of audio in a single pass.

Photo of author

Sophie Lin - Technology Editor

Sophie is a tech innovator and acclaimed tech writer recognized by the Online News Association. She translates the fast-paced world of technology, AI, and digital trends into compelling stories for readers of all backgrounds.

AEW Reportedly Discussed Six-Man Match Similar To WWE’s Elimination Chamber

US Troops Withdrawal from Iraq Marks End of American Hegemony in Region