Google DeepMind has released EmbeddingGemma 2, an open-source multimodal embedding model that maps text, code, images, video, and audio into a unified 768-dimensional vector space. The model consists of 740 million total parameters, combining a 270M parameter text encoder with modular vision (170M) and audio (300M) encoders.

  • Designed to run on consumer hardware like mobile devices and laptops for low-latency semantic representations.
  • Supports on-device applications including search, retrieval-augmented generation (RAG), classification, and clustering.
  • Available as an open model via Hugging Face.