Google has released EmbeddingGemma 2, an open-source model designed for natively multimodal embeddings.
The release positions the new model as a best-in-class option for handling both text and image data within a unified embedding space.
Google has released EmbeddingGemma 2, an open-source model designed for natively multimodal embeddings.
The release positions the new model as a best-in-class option for handling both text and image data within a unified embedding space.
Google DeepMind has released EmbeddingGemma 2, an open-source multimodal embedding model that maps text, code, images, video, and audio into a unified 768-dimensional vector space. The model consists of 740 million total parameters, combining a 270M parameter text encoder with modular vision (170M) and audio (300M) encoders.
Google introduces Gemma 4, a new generation of open-weight, natively multimodal language models featuring dense and Mixture-of-Experts architectures ranging from 2.3B to 31B parameters.
Google introduces Gemma 3, a new family of lightweight open multimodal models ranging from 1 to 27 billion parameters. This update adds vision understanding, support for at least 128K tokens of context, and expanded multilingual coverage compared to previous versions.
Researchers introduce Eduardo, a multi-turn reinforcement learning recipe for training large language model tutors that addresses the "assistance dilemma" where models tend to give answers rather than teach. The method uses a masked near-transfer post-test and binary reward gates to discourage cognitive offloading and prevent solution handover.
The authors introduce Index-Translate, a multilingual translation model family combining a shared foundation with specialized training for general translation, instruction following, speech translation, controlled dubbing, and long-document translation. The family includes 2B, 9B, and 35B-A3B models supporting 150 languages.
We use cookies to measure traffic and improve the site. You can accept or decline analytics cookies. Privacy policy