Google and Google DeepMind have introduced Gemini 1.0, a new family of large-scale multimodal models designed to understand and combine text, code, audio, image, and video natively.
- The first version includes three optimized sizes: Gemini Ultra for complex tasks, Gemini Pro for scaling across tasks, and Gemini Nano for on-device efficiency.
- Gemini Ultra achieves state-of-the-art performance on 30 of 32 widely-used academic benchmarks and is the first model to outperform human experts on MMLU with a score of 90.0%.
- The models are built from the ground up to be multimodal, allowing them to reason about diverse inputs without relying on separate components stitched together.
This release represents a significant step in building general-purpose AI that can enhance how developers and enterprises build and scale applications.