Benchmark · math

AIME 2025

31 results 27 models

AIME 2025 tests a model's advanced mathematical reasoning using the 15 problems from the 2025 American Invitational Mathematics Examination, a hard high-school competition. Each answer is an integer from 0 to 999, and a model's score is the percentage of problems solved correctly (often reported as pass@1 or avg@k).

Read more
Example
A typical item is a challenging competition problem — for instance, a number-theory or combinatorics question whose solution is a single integer between 0 and 999, such as finding the number of ordered pairs that satisfy a given set of algebraic conditions.
Scoring
Each of the 15 problems is graded by exact match on its integer answer, and the score is the fraction (percentage) solved correctly; because answers vary from run to run, results are usually averaged over many samples (avg@k) or reported as pass@1.
Verification
Verification is automatic and objective: the model's final integer is compared to the official answer key with no partial credit and no human judging, so a solution is accepted only on an exact numeric match.
Why it matters
AIME problems demand multi-step, creative reasoning rather than recall, so the benchmark is a widely watched measure of frontier LLM math ability and a common headline number when new reasoning models are released.
Worked example
Task
Representative AIME-style problem (answer is an integer from 0 to 999): Find the sum of all positive integers n for which √(n² + 85n + 2017) is an integer.
Solution
Complete the square: (2m)² − (2n+85)² = 843 = 3·281. Factor the difference of squares: (2m−2n−85)(2m+2n+85) = 843, giving n = 168 or n = 27. Sum = 168 + 27 = 195.
Walkthrough
Completing the square turns the condition into a difference of squares (2m)² − (2n+85)² = 843 = 3·281, whose only positive factorizations give n = 168 and n = 27, summing to 195. AIME is graded by exact match of the single integer answer (0–999), with no partial credit.
0 25 50 75 100 2025-01-22 2025-10-17 2026-07-13 Kimi k1.5 (short-CoT) · 60.8 · 2025-01-22 Kimi k1.5 · 77.5 · 2025-01-22 Grok 3 (Think) · 93.3 · 2025-02-19 Grok 3 Beta · 93.3 · 2025-02-26 OpenAI o3-mini · 83.3 · 2025-02-26 DeepSeek R1 · 79.8 · 2025-02-26 Claude 3.7 Sonnet · 61.3 · 2025-02-26 Claude 3.5 Sonnet · 41.3 · 2025-02-26 Grok 3 · 93.3 · 2025-03-11 Gemini 2.5 Pro (experimental) · 86.7 · 2025-03-25 Gemini 2.5 Pro · 86.7 · 2025-03-26 OpenAI o3 · 98.4 · 2025-04-16 OpenAI o4-mini · 99.5 · 2025-04-16 o4-mini · 92.7 · 2025-04-17 Qwen3-235B-A22B · 81.5 · 2025-05-14 DeepSeek-R1-0528 · 87.5 · 2025-05-28 Deepseek-R1-0528 · 87.5 · 2025-05-30 Grok 4 · 88 · 2025-07-25 Kimi K2 · 49.5 · 2025-07-28 Kimi K2 · 49.5 · 2025-07-28 GPT-5 · 94.6 · 2025-08-07 GPT-5 · 94.6 · 2025-08-07 GPT-5 · 94.6 · 2025-08-07 GLM-4.6 · 93.9 · 2025-09-30 Claude Sonnet 4.5 · 100 · 2025-09-30 Kimi K2.5 · 96.1 · 2026-01-27 o1-pro · 86 · 2026-05-14 Agents-A1 · 79 · 2026-06-30 DASH · 50.8 · 2026-07-05 Mach-Mind-4-Flash · 92.7 · 2026-07-13 Mach-Mind-4-Flash · 92.7 · 2026-07-13
Kimi k1.5 (short-CoT) Kimi k1.5 Grok 3 (Think) Grok 3 Beta OpenAI o3-mini DeepSeek R1 Claude 3.7 Sonnet Claude 3.5 Sonnet Grok 3 Gemini 2.5 Pro (experimental) Gemini 2.5 Pro OpenAI o3 OpenAI o4-mini o4-mini Qwen3-235B-A22B DeepSeek-R1-0528 Deepseek-R1-0528 Grok 4 Kimi K2 GPT-5 GLM-4.6 Claude Sonnet 4.5 Kimi K2.5 o1-pro Agents-A1 DASH Mach-Mind-4-Flash
Timeline
Date Model Score Source
2026-07-13 Mach-Mind-4-Flash 92.7% Mach-Mind-4-Flash: 35B MoE agentic model matches larger models via post-training optimization
2026-07-13 Mach-Mind-4-Flash 92.7% Mach-Mind-4-Flash: 35B MoE agentic model matches larger models via post-training optimization
2026-07-05 DASH 50.8% DASH reduces overthinking in reasoning language models via segment-level credit assignment
2026-06-30 Agents-A1 79.0pts Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent
2026-05-14 o1-pro 86.0% OpenAI releases o1 family establishing the reasoning era with inference-time compute
2026-01-27 Kimi K2.5 96.1% Moonshot releases Kimi K2.5 with Agent Swarm technology
2025-09-30 GLM-4.6 93.9% Zhipu AI releases GLM-4.6 with agentic capabilities and 128k context window
2025-09-30 Claude Sonnet 4.5 100.0% Anthropic releases Claude Sonnet 4.5 with improved coding and agent capabilities
2025-08-07 GPT-5 94.6% OpenAI launches GPT-5 with adaptive reasoning and unified architecture
2025-08-07 GPT-5 94.6% OpenAI introduces GPT-5 with unified routing and expert-level reasoning
2025-08-07 GPT-5 94.6% OpenAI
2025-07-28 Kimi K2 49.5% Kimi K2: open MoE model with 1T parameters and MuonClip optimizer
2025-07-28 Kimi K2 49.5% Moonshot releases Kimi K2, an open agentic MoE model with 32B activated parameters
2025-07-25 Grok 4 88.0% Epoch AI evaluates Grok 4's math capabilities
2025-05-30 Deepseek-R1-0528 87.5% DeepSeek updates R1 model with improved reasoning and benchmark scores
2025-05-28 DeepSeek-R1-0528 87.5% DeepSeek-R1-0528 improves reasoning and AIME accuracy to 87.5%
2025-05-14 Qwen3-235B-A22B 81.5% Qwen3 introduces unified thinking mode and 119-language support
2025-04-17 o4-mini 92.7% OpenAI releases o4-mini, a faster, cheaper multimodal reasoning model
2025-04-16 OpenAI o3 98.4% OpenAI releases o3 and o4-mini reasoning models with agentic tool use
2025-04-16 OpenAI o4-mini 99.5% OpenAI releases o3 and o4-mini reasoning models with agentic tool use
2025-03-26 Gemini 2.5 Pro 86.7% Google releases experimental Gemini 2.5 Pro reasoning model with 1M token context
2025-03-25 Gemini 2.5 Pro (experimental) 86.7% Google DeepMind releases Gemini 2.5 Pro experimental model
2025-03-11 Grok 3 93.3% xAI releases Grok 3 with deep reasoning and million-token context
2025-02-26 Grok 3 Beta 93.3% Benchmark comparison of Claude 3.7 Sonnet, o3-mini, R1, and Grok 3 Beta
2025-02-26 OpenAI o3-mini 83.3% Benchmark comparison of Claude 3.7 Sonnet, o3-mini, R1, and Grok 3 Beta
2025-02-26 DeepSeek R1 79.8% Benchmark comparison of Claude 3.7 Sonnet, o3-mini, R1, and Grok 3 Beta
2025-02-26 Claude 3.7 Sonnet 61.3% Benchmark comparison of Claude 3.7 Sonnet, o3-mini, R1, and Grok 3 Beta
2025-02-26 Claude 3.5 Sonnet 41.3% Benchmark comparison of Claude 3.7 Sonnet, o3-mini, R1, and Grok 3 Beta
2025-02-19 Grok 3 (Think) 93.3% xAI releases Grok 3 Beta and DeepSearch agent
2025-01-22 Kimi k1.5 (short-CoT) 60.8% Kimi k1.5 scales reinforcement learning with LLMs to match o1 and beat short-CoT models
2025-01-22 Kimi k1.5 77.5% Kimi k1.5 scales reinforcement learning with LLMs to match o1 and beat short-CoT models