Skip to content
Apixo
Blog
news· 2 min read· via SiliconANGLE AI

Google Releases EmbeddingGemma 2 With Multimodal Capabilities

Google has launched EmbeddingGemma 2, an open multimodal embedding model capable of running locally on smartphones and handling text, images, audio, and video.

Google Releases EmbeddingGemma 2 With Multimodal Capabilities

Google LLC has expanded its lightweight embedding technology by releasing EmbeddingGemma 2, an open multimodal model designed to operate directly on mobile devices like smartphones. While the initial iteration released in September 2025 focused exclusively on text, the updated version integrates text, images, audio, and video into a single unified embedding space.

According to Google DeepMind research engineers Sahil Dua and Henrique Schechter Vera, reception to the initial release exceeded projections, accumulating more than 20 million downloads by developers.

Architecture and Performance Updates

Constructed on top of the Gemma 4 architecture introduced in April, EmbeddingGemma 2 measures 740 million parameters, more than doubling the scale of its predecessor. The majority of this expansion supports the vision and audio encoders, which text-only applications can omit entirely. During testing of a quantized build on a Google Pixel 11 Pro, the 270 million-parameter text core consumed approximately 191 megabytes of memory independently.

Applications utilizing the model must account for storage of the generated outputs, where each embedding comprises a list of 768 numbers added to a local vector database for every indexed photo, clip, or document. To mitigate storage overhead, the model incorporates Matryoshka Representation Learning, allowing developers to truncate these lists down to 128 numbers to reduce space requirements by up to six times. Documentation indicates that retaining 256 numbers preserves roughly 95% of full retrieval quality across image, video, and speech tasks.

Performance metrics highlight a notable improvement in code handling, with EmbeddingGemma 2 achieving 78.68 on the code section of the Massive Text Embedding Benchmark, nearly 10 points higher than the original version. Google aims this capability at developers engineering retrieval systems for coding agents, whereas multilingual text performance remained largely steady. Furthermore, Google positions the model as a top performer among multimodal embedding options under 1 billion parameters, outperforming certain specialized models more than twice its size across visual and auditory benchmarks.

What it means for developers

For developers building on-device applications, the shared text tokenizer and audio encoder between EmbeddingGemma 2 and Gemma 4 reduce memory footprints when implementing retrieval-augmented generation setups compared to running separate models. Implementations are already emerging, such as Google’s AI Edge Foresight meeting app for Mac operating both models concurrently, and an AI Edge Gallery demo featuring a Video Moments Finder to locate specific scenes using typed or spoken prompts.

Developers looking to experiment with a wide range of top foundational models can try top AI models cheaply through one API at https://apixoai.online.

The model weights are currently accessible via Hugging Face and Kaggle under an Apache 2.0 license permitting commercial use, with upcoming integration planned for the Model Garden within the Gemini Enterprise Agent Platform. The weights are also compatible with open-source serving infrastructure including vLLM, llama.cpp, and Ollama.


Source: Google expands EmbeddingGemma beyond text to images, audio and video — SiliconANGLE AI. Written by the Apixo team from that report.

#ai-news#google#embeddinggemma#ai#multimodal#developers
Try it with your own tools

One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.

Get your API key

Keep reading