Benchmark · math

AIME 2024

30 results 27 models

AIME 2024 measures a model's advanced mathematical reasoning using the 15 problems from the 2024 American Invitational Mathematics Examination, reporting the percentage solved correctly.

Read more
Example
A typical item is a hard high-school competition problem — for example a number-theory or combinatorics question whose answer is a single integer (each AIME answer is an integer, conventionally from 0 to 999).
Scoring
The metric is accuracy: the share of the 15 problems answered correctly, so each solved problem is worth 1/15 of the score.
Verification
A result is accepted by exact match — the model's final integer answer must equal the official answer, checked automatically with no partial credit and no human judging.
Why it matters
Because it is a compact set of hard, recent competition problems that demand multi-step reasoning, it is widely used to compare frontier models, and its recency limits training-data contamination.
Worked example
Task
How many positive integers n ≤ 600 are divisible by 4 or by 6 but not by 5? (AIME answers are integers from 0 to 999.)
Solution
Multiples of 4 or 6 that are ≤ 600: 150 + 100 − 50 = 200. Of these, the ones also divisible by 5 (multiples of 20 or 30, minus 60): 30 + 20 − 10 = 40. Answer: 200 − 40 = 160.
Walkthrough
By inclusion–exclusion, 200 integers are divisible by 4 or 6, and 40 of those are also divisible by 5, leaving 160. AIME auto-grades only the final integer (0–999), with no partial credit for the work.
0 25 50 75 100 2024-09-12 2025-08-09 2026-07-07 o1 · 13.9 · 2024-09-12 OpenAI o1-preview · 83 · 2024-10-01 QwQ-32B-Preview · 50 · 2024-11-28 o3 · 96.7 · 2024-12-20 o3 · 96.7 · 2026-03-25 DeepSeek-R1-Distill-Qwen-7B · 55.5 · 2025-01-22 DeepSeek-R1-Zero · 71 · 2025-01-22 DeepSeek-R1-Distill-Qwen-32B · 72.6 · 2025-01-22 DeepSeek-R1 · 79.8 · 2025-01-22 s1-32B · 57 · 2025-01-31 s1-32B · 57 · 2025-01-31 LIMO · 57.1 · 2025-02-05 Grok 3 mini · 95.8 · 2025-02-19 Claude 3.7 Sonnet · 80 · 2025-02-24 GPT-4.5 · 36.7 · 2025-03-05 DeepSeek-V3-0324 · 59.4 · 2025-03-24 Gemini 2.5 Pro · 92 · 2025-03-26 Seed1.5-Thinking · 86.7 · 2025-04-10 o4-mini · 93.4 · 2025-04-17 Qwen3-235B-A22B · 85.7 · 2025-05-14 Deepseek-R1-0528-Qwen3-8B · 86 · 2025-05-30 Deepseek-R1-0528 · 91.4 · 2025-05-30 Magistral Small · 70.7 · 2025-06-10 Magistral Medium · 73.6 · 2025-06-10 Magistral Medium · 50 · 2025-06-12 GLM-4.5 · 91 · 2025-08-06 GLM-4.5-Air · 89.4 · 2025-08-06 Qwen3-1.7B · 62.4 · 2026-07-07 DeepSeek R1 · 79.8 · 2025-01-20 OpenAI o3 · 91.6 · 2025-04-16
o1 OpenAI o1-preview QwQ-32B-Preview o3 DeepSeek-R1-Distill-Qwen-7B DeepSeek-R1-Zero DeepSeek-R1-Distill-Qwen-32B DeepSeek-R1 s1-32B LIMO Grok 3 mini Claude 3.7 Sonnet GPT-4.5 DeepSeek-V3-0324 Gemini 2.5 Pro Seed1.5-Thinking o4-mini Qwen3-235B-A22B Deepseek-R1-0528-Qwen3-8B Deepseek-R1-0528 Magistral Small Magistral Medium GLM-4.5 GLM-4.5-Air Qwen3-1.7B DeepSeek R1 OpenAI o3
Timeline
Date Model Score Source
2026-07-07 Qwen3-1.7B 62.4% Direct-OPD distills weak-to-strong RL gains for stronger models
2026-03-25 o3 96.7% OpenAI announces o3 retirement from ChatGPT and shares benchmark scores
2025-08-06 GLM-4.5 91.0% Zhipu AI releases GLM-4.5 models matching Claude and DeepSeek
2025-08-06 GLM-4.5-Air 89.4% Zhipu AI releases GLM-4.5 models matching Claude and DeepSeek
2025-06-12 Magistral Medium 50.0% Mistral introduces Magistral reasoning models with scalable RL pipeline
2025-06-10 Magistral Small 70.7% Mistral AI releases Magistral reasoning model with open and enterprise variants
2025-06-10 Magistral Medium 73.6% Mistral AI releases Magistral reasoning model with open and enterprise variants
2025-05-30 Deepseek-R1-0528-Qwen3-8B 86.0% DeepSeek updates R1 model with improved reasoning and benchmark scores
2025-05-30 Deepseek-R1-0528 91.4% DeepSeek updates R1 model with improved reasoning and benchmark scores
2025-05-14 Qwen3-235B-A22B 85.7% Qwen3 introduces unified thinking mode and 119-language support
2025-04-17 o4-mini 93.4% OpenAI releases o4-mini, a faster, cheaper multimodal reasoning model
2025-04-16 OpenAI o3 91.6% OpenAI
2025-04-10 Seed1.5-Thinking 86.7% Seed1.5-Thinking uses reinforcement learning to improve reasoning
2025-03-26 Gemini 2.5 Pro 92.0% Google releases experimental Gemini 2.5 Pro reasoning model with 1M token context
2025-03-24 DeepSeek-V3-0324 59.4% DeepSeek-V3-0324 improves reasoning, coding, and Chinese writing over DeepSeek-V3
2025-03-05 GPT-4.5 36.7% OpenAI releases GPT-4.5, its largest model, for ChatGPT Plus
2025-02-24 Claude 3.7 Sonnet 80.0% Anthropic releases Claude 3.7 Sonnet with dynamic reasoning control
2025-02-19 Grok 3 mini 95.8% xAI releases Grok 3 Beta and DeepSearch agent
2025-02-05 LIMO 57.1% GAIR-NLP releases LIMO, achieving strong math reasoning with 817 samples
2025-01-31 s1-32B 57.0% Simple Scaling Lab releases s1-32B with test-time scaling via budget forcing
2025-01-31 s1-32B 57.0% s1-32B exceeds o1-preview on competition math using simple test-time scaling
2025-01-22 DeepSeek-R1-Distill-Qwen-7B 55.5% DeepSeek releases R1 reasoning models via reinforcement learning
2025-01-22 DeepSeek-R1-Zero 71.0% DeepSeek releases R1 reasoning models via reinforcement learning
2025-01-22 DeepSeek-R1-Distill-Qwen-32B 72.6% DeepSeek releases R1 reasoning models via reinforcement learning
2025-01-22 DeepSeek-R1 79.8% DeepSeek releases R1 reasoning models via reinforcement learning
2025-01-20 DeepSeek R1 79.8% DeepSeek
2024-12-20 o3 96.7% OpenAI releases o3 model with high performance in reasoning and coding
2024-11-28 QwQ-32B-Preview 50.0% Qwen releases QwQ-32B-Preview, an experimental reasoning model
2024-10-01 OpenAI o1-preview 83.0% Comparing OpenAI o1 to other Top Models
2024-09-12 o1 13.9% OpenAI releases o1-preview, a reasoning model that surpasses human experts on math and science benchmarks