Google introduces Gemma 3, a new family of lightweight open multimodal models ranging from 1 to 27 billion parameters. This update adds vision understanding, support for at least 128K tokens of context, and expanded multilingual coverage compared to previous versions.
- Models utilize a tailored SigLIP vision encoder with a Pan & Scan method for flexible image resolution.
- Architecture changes include a 5:1 interleaving of local to global attention layers to reduce KV-cache memory usage during long-context inference.
- The 1B model supports 32K context, while larger models support 128K tokens.
- Training employs knowledge distillation from Gemini frontier models and a novel post-training recipe.
- Gemma3-4B-IT matches Gemma2-27B-IT performance, and Gemma3-27B-IT is comparable to Gemini-1.5-Pro on benchmarks.
The release provides all model variants to the community, enabling powerful multimodal capabilities on standard consumer-grade hardware.