Google has introduced Gemini 2.5, its newest family of "thinking" AI models capable of reasoning through thoughts before responding. The first release is an experimental version of Gemini 2.5 Pro, which achieves state-of-the-art performance on a wide range of benchmarks and debuts at #1 on the LMArena leaderboard.
- Gemini 2.5 Pro tops the LMArena leaderboard by a significant margin, indicating high-quality style and capability.
- It leads on math and science benchmarks like GPQA and AIME 2025 without using test-time techniques like majority voting.
- The model scores a state-of-the-art 18.8% on Humanity’s Last Exam, a dataset capturing the human frontier of knowledge.
- On SWE-Bench Verified, Gemini 2.5 Pro scores 63.8% with a custom agent setup for agentic code evaluation.
- It features a 1 million token context window, with 2 million tokens coming soon, supporting text, audio, images, video, and code repositories.
Gemini 2.5 Pro is available now in Google AI Studio and the Gemini app for Advanced users, with Vertex AI availability expected in the coming weeks.