Benchmark · math

FrontierMath

20 results 15 models

FrontierMath is a benchmark from Epoch AI that measures how well an AI model can solve extremely hard, original research-level mathematics problems. The score is the percentage of problems the model answers correctly, and today's models solve only a small fraction — far from saturation.

Read more
Example
A single, previously unpublished problem from an advanced field such as number theory, algebraic geometry, or combinatorics, whose solution demands deep expertise and reduces to one definite final answer (for example a specific integer or an exact mathematical object).
Scoring
The metric is accuracy: the fraction of problems whose final answer exactly matches the reference answer, reported as a percentage.
Verification
Each problem has a single definite, automatically checkable answer, so a solution is accepted only when it matches the reference exactly; the problems are crafted by expert mathematicians to be almost impossible to get right by guessing.
Why it matters
Because the problems are novel it resists memorization, and it remains far from being solved, making it one of the few math benchmarks that still cleanly separates genuine advanced mathematical reasoning from pattern-matching.
Worked example
Task
Determine the number of integers $n$ with $0 \le n < 3^{13}$ for which the central binomial coefficient $\binom{2n}{n}$ is not divisible by $3$.
Solution
By Kummer's theorem, $3 \nmid \binom{2n}{n}$ exactly when adding $n+n$ in base 3 produces no carries, i.e. every base-3 digit of $n$ is 0 or 1. Over $0 \le n < 3^{13}$ this gives 2 choices for each of the 13 digits, so the count is $2^{13} = 8192$.
Walkthrough
Kummer's theorem gives $v_3\binom{2n}{n}$ as the number of carries when adding $n$ to itself in base 3; zero carries forces each base-3 digit to be 0 or 1, i.e. 2 choices per digit over 13 digits. Grading: an automated checker exact-matches the single final integer (8192).
0 14 28 42 56 2024-11-08 2025-08-01 2026-04-25 Claude 3.5 Sonnet · 2 · 2024-11-08 Gemini 1.5 Pro · 2 · 2024-11-08 GPT-4o · 2 · 2024-11-08 o1-preview · 2 · 2024-11-08 o3 · 25 · 2024-12-20 o3 · 25 · 2024-12-22 o3 · 25.2 · 2025-01-19 o3 · 25 · 2025-01-29 o3 · 10 · 2025-04-20 o3-mini with high reasoning and a Python tool · 32 · 2025-03-17 Gemini 2.5 Deep Think · 29 · 2025-10-09 GPT-5 Pro · 13 · 2025-10-13 Claude 4.5 Opus · 21 · 2025-12-04 Gemini 3 Pro · 38 · 2025-12-04 GPT-5.2 Thinking · 40.3 · 2025-12-11 GPT-5.2 Pro · 31 · 2026-01-23 GPT-5.5 Pro · 39.6 · 2026-04-23 GPT-5.5 · 35.4 · 2026-04-23 GPT-5.5 · 51.7 · 2026-04-25 OpenAI o3 · 25.2 · 2024-12-20
Claude 3.5 Sonnet Gemini 1.5 Pro GPT-4o o1-preview o3 o3-mini with high reasoning and a Python tool Gemini 2.5 Deep Think GPT-5 Pro Claude 4.5 Opus Gemini 3 Pro GPT-5.2 Thinking GPT-5.2 Pro GPT-5.5 Pro GPT-5.5 OpenAI o3
Timeline
Date Model Score Source
2026-04-25 GPT-5.5 51.7% OpenAI releases GPT-5.5, a ground-up retrained model with native omnimodal processing
2026-04-23 GPT-5.5 Pro 39.6% OpenAI releases GPT-5.5 with agentic capabilities and doubled API pricing
2026-04-23 GPT-5.5 35.4% OpenAI releases GPT-5.5 with agentic capabilities and doubled API pricing
2026-01-23 GPT-5.2 Pro 31.0% GPT-5.2 Pro scores 31% on FrontierMath Tier 4
2025-12-11 GPT-5.2 Thinking 40.3% OpenAI releases GPT-5.2 Pro and Thinking models for science and math
2025-12-04 Claude 4.5 Opus 21.0% Epoch AI launches Frontier Data Centers Hub and analyzes OSWorld benchmark
2025-12-04 Gemini 3 Pro 38.0% Epoch AI launches Frontier Data Centers Hub and analyzes OSWorld benchmark
2025-10-13 GPT-5 Pro 13.0% GPT-5 Pro solves new FrontierMath Tier 4 problem, edging Gemini 2.5 Deep Think
2025-10-09 Gemini 2.5 Deep Think 29.0% Epoch AI evaluates Gemini 2.5 Deep Think's math capabilities
2025-04-20 o3 10.0% OpenAI's o3 scores lower on FrontierMath than initially claimed
2025-03-17 o3-mini with high reasoning and a Python tool 32.0% Epoch's evaluation of o3-mini on FrontierMath yields 11% vs OpenAI's 32%
2025-01-29 o3 25.0% OpenAI's o3 model achieves 25% score on Frontier Math benchmark
2025-01-19 o3 25.2% OpenAI funded FrontierMath benchmark before o3 set record
2024-12-22 o3 25.0% OpenAI's o3 scores 25% on FrontierMath hard math benchmark
2024-12-20 o3 25.0% OpenAI previews o3 reasoning model with ARC AGI and Frontier Math breakthroughs
2024-12-20 OpenAI o3 25.2% Epoch AI
2024-11-08 Claude 3.5 Sonnet 2.0% Epoch AI releases FrontierMath benchmark to evaluate advanced mathematical reasoning
2024-11-08 Gemini 1.5 Pro 2.0% Epoch AI releases FrontierMath benchmark to evaluate advanced mathematical reasoning
2024-11-08 GPT-4o 2.0% Epoch AI releases FrontierMath benchmark to evaluate advanced mathematical reasoning
2024-11-08 o1-preview 2.0% Epoch AI releases FrontierMath benchmark to evaluate advanced mathematical reasoning