Google DeepMind has released EmbeddingGemma 2, an open-source multimodal embedding model built on the Gemma 4 architecture that embeds text, code, images, video, and audio into a shared 768-dimensional vector space. The model features 740 million parameters with a modular design allowing for smaller text-only or vision-audio variants, and it supports an 8K token context window.
- EmbeddingGemma 2 achieves leading scores among sub-1B multimodal embedders, including 78.68 on MTEB Code v1 and 49.39 on MAEB (audio).
- The model is optimized for on-device deployment, requiring approximately 191MB of RAM for text-only weights and 567MB for the full multimodal model on a Pixel 11 Pro.
- It utilizes Matryoshka Representation Learning to allow vector truncation to 512, 256, or 128 dimensions, reducing storage requirements significantly with minimal performance loss.
- Weights are available under an Apache 2.0 license on Hugging Face and Kaggle, with support for Ollama, llama.cpp, LiteRT, and other frameworks.
The release enables privacy-first retrieval-augmented generation (RAG) and local search capabilities by allowing developers to run embeddings directly on devices without sending data to the cloud.