Google introduces Gemma 3, a new family of lightweight open multimodal models ranging from 1 to 27 billion parameters. This update adds vision understanding, support for at least 128K tokens of context, and expanded multilingual coverage compared to previous versions.

  • Models utilize a tailored SigLIP vision encoder with a Pan & Scan method for flexible image resolution.
  • Architecture changes include a 5:1 interleaving of local to global attention layers to reduce KV-cache memory usage during long-context inference.
  • The 1B model supports 32K context, while larger models support 128K tokens.
  • Training employs knowledge distillation from Gemini frontier models and a novel post-training recipe.
  • Gemma3-4B-IT matches Gemma2-27B-IT performance, and Gemma3-27B-IT is comparable to Gemini-1.5-Pro on benchmarks.

The release provides all model variants to the community, enabling powerful multimodal capabilities on standard consumer-grade hardware.