Skip to content
Apixo
Blog
news· 2 min read· via Google DeepMind

Google DeepMind Releases EmbeddingGemma 2 for On-Device Multimodal AI

Google DeepMind has launched EmbeddingGemma 2, a 740M-parameter multimodal embedding model that natively maps text, images, audio, and video into a unified vector space.

Google DeepMind Releases EmbeddingGemma 2 for On-Device Multimodal AI

Google DeepMind has introduced EmbeddingGemma 2, building upon the original text-only release that achieved over 20 million downloads for local search tools and privacy-first retrieval-augmented generation pipelines. The newly released model expands capabilities to natively map combinations of text, code, images, audio, and video into a unified embedding space on consumer hardware.

Constructed on the Gemma 4 architecture and distributed under a commercially permissive Apache 2.0 license, EmbeddingGemma 2 features 740 million parameters designed specifically for on-device inference. Research engineers Sahil Dua and Henrique Schechter Vera noted that the system can locate a specific video clip from a voice memo or search through hours of audio records using a text query within a single natively multimodal framework.

Performance and Architecture Specs

EmbeddingGemma 2 is engineered to deliver high performance within strict resource boundaries:

  • Size and Benchmarks: It achieves leading scores among sub-1 billion multimodal embedders on benchmarks such as MTEB, MTEB Code, and MAEB. Compared to its predecessor, it delivers a 9.92-point improvement in code performance on MTEB Code, moving from 68.76 to 78.68.
  • Modular Design: The base model requires 270 million parameters for text-only workflows, with optional vision (170M) and audio (300M) encoders added for complete multimodal support.
  • Storage Efficiency: Utilizing Matryoshka Representation Learning (MRL), output vectors can be dynamically scaled down from 768 dimensions to 512, 256, or 128 dimensions, providing up to a 6x reduction in storage requirements for local vector databases.
  • Hardware Optimization: When quantized, the model operates efficiently on devices like the Google Pixel 11 Pro, consuming roughly 191MB of active RAM for text-only weights and about 567MB for the full multimodal configuration.
  • Extended Context: The context window is expanded to 8K tokens—four times larger than the first version—allowing the local processing of up to 5.5 minutes of audio, 29 images, 58 video frames, or interleaved combinations.

What it means for developers

For engineers building privacy-first applications, local embeddings ensure data privacy, decrease pipeline latency, and enable offline cross-modal search and retrieval. Because EmbeddingGemma 2 shares its text tokenizer and audio encoder with Gemma 4, both systems can run concurrently in a unified pipeline with a reduced combined memory footprint. Developers looking to experiment broadly can also try top AI models cheaply through one API at https://apixoai.online.

Deployment options span multiple ecosystems. Weights are available on Hugging Face and Kaggle, with upcoming availability on the Gemini Enterprise Agent Platform Model Garden. Developers can use Google AI Edge MediaPipe for turnkey tasks, LiteRT for custom integration, or run models in the browser via transformers.js and WebGPU. Serving tools like vLLM, Ollama, MLX, llama.cpp, SGLang, and LMStudio are supported, alongside vector storage integration with Qdrant and fine-tuning guides provided by Unsloth.


Source: EmbeddingGemma 2: an open, lightweight multimodal embedding model — Google DeepMind. Written by the Apixo team from that report.

#ai-news#google#gemma#embeddings#multimodal#ai
Try it with your own tools

One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.

Get your API key

Keep reading