Benchmark · reasoning

GPQA Diamond

61 results 44 models

GPQA Diamond is the hardest split of GPQA, a set of graduate-level, "Google-proof" multiple-choice science questions in biology, physics, and chemistry written by domain experts. A model is scored by its percentage accuracy at picking the correct answer.

Read more
Example
A typical item is a hard multiple-choice question that needs graduate-level reasoning — for instance, a chemistry question describing a reaction mechanism and asking which of four options gives the correct product, so tough that a non-expert can't solve it even with web search.
Scoring
The score is the percentage of questions answered correctly (accuracy): the number of items where the model picks the right option divided by the total number of questions.
Verification
Each answer is graded automatically by exact match against the single known-correct option, so no human judging is needed; the "Diamond" label itself was assigned during dataset creation, when qualified experts agreed on the answer and non-experts still got it wrong.
Why it matters
It is one of the toughest tests of genuine expert-level scientific reasoning (not facts you can look up), so a strong GPQA Diamond score is a headline signal that a frontier model can reason through hard, specialized problems.
Worked example
Task
An electron is in the (unnormalized) spin state (3i, 4)ᵀ. Find the expectation value of its spin along the y-axis, S_y. Options: (A) 12ħ/25 (B) −12ħ/25 (C) 25ħ/2 (D) −25ħ/2.
Solution
Norm: |3i|²+|4|²=25. σ_y=[[0,−i],[i,0]]. ⟨S_y⟩=(ħ/2)·(ψ†σ_yψ)/(ψ†ψ)=(ħ/2)·(−24/25)=−12ħ/25 → (B).
Walkthrough
Using ⟨S_y⟩ = ψ†S_yψ / ψ†ψ with S_y = (ħ/2)σ_y, the numerator ψ†σ_yψ = −24 and the norm ψ†ψ = 25, giving −12ħ/25 (option B). GPQA grades by exact match of the selected letter against the gold key.
0 25 50 75 100 2023-11-20 2025-03-27 2026-08-03 GPT-4 based baseline · 39 · 2023-11-20 Claude 3.5 Sonnet · 59.4 · 2024-06-20 Claude 3.5 Sonnet · 59.4 · 2024-06-20 Claude 3.5 Sonnet · 65 · 2025-02-26 Qwen2-72B · 37.9 · 2024-07-15 Qwen2-72B · 37.9 · 2024-07-15 QwQ-32B-Preview · 65.2 · 2024-11-28 o3 · 87.7 · 2024-12-20 o3 · 87.7 · 2026-03-25 Grok 3 (Think) · 84.6 · 2025-02-19 Claude 3.7 Sonnet · 84.8 · 2025-02-24 Claude 3.7 Sonnet · 84.8 · 2025-02-26 Grok 3 Beta · 84.6 · 2025-02-26 DeepSeek R1 · 71.5 · 2025-02-26 OpenAI o3-mini · 78 · 2025-02-26 GPT-4.5 · 71.4 · 2025-03-05 Grok 3 · 84.6 · 2025-03-11 DeepSeek-V3-0324 · 68.4 · 2025-03-24 Gemini 2.5 Pro (experimental) · 84 · 2025-03-25 Gemini 2.5 Pro · 84 · 2025-03-25 Gemini 2.5 Pro · 84 · 2025-03-26 Gemini 2.5 Pro · 84 · 2025-04-05 Seed1.5-Thinking · 77.3 · 2025-04-10 GPT‑4.1 nano · 50.3 · 2025-04-14 Gemini 2.0 Flash · 60.1 · 2025-04-15 Claude Sonnet 4 · 70 · 2025-05-22 Claude Opus 4 · 74.9 · 2025-05-22 Claude Opus 4 · 74.9 · 2025-05-22 Deepseek-R1-0528 · 81 · 2025-05-30 Kimi K2 · 75.1 · 2025-07-28 Kimi K2 · 75.1 · 2025-07-28 GPT-5 pro · 88.4 · 2025-08-07 GPT-5 pro · 89.4 · 2025-08-07 GPT-5 Pro · 88.4 · 2025-08-07 DeepSeek-V3.1-Terminus · 74.2 · 2025-09-22 Claude Sonnet 4.5 · 83.4 · 2025-09-30 GPT-5.2 · 92.4 · 2025-10-24 GPT-4 · 39 · 2025-10-24 Claude 3 Opus · 60 · 2025-10-24 O1 · 77 · 2025-10-24 Claude Opus 4.6 · 91.3 · 2025-10-24 Claude Opus 4.6 · 91.3 · 2026-02-05 Gemini 3 Pro · 91.9 · 2025-10-24 Gemini 3 Pro · 91.9 · 2025-11-18 Gemini 3 Pro · 91.9 · 2025-11-18 Gemini 3 Pro · 93 · 2025-12-04 Gemini 3.1 Pro Preview · 94.1 · 2025-10-24 Aristotle-X1 · 92.4 · 2025-10-24 Gemini 3 Deep Think · 93.8 · 2025-11-18 Gemini 3 Deep Think · 93.8 · 2025-11-18 GPT-5.2 Thinking · 92.4 · 2025-12-11 GPT-5.2 Thinking · 92.4 · 2025-12-11 GPT-5.2 Pro · 93.2 · 2025-12-11 GPT-5.2 Pro · 93.2 · 2025-12-11 Kimi K2.5 · 87.6 · 2026-01-27 Grok 4 · 87.7 · 2026-02-25 GPT-Live-1 · 84.2 · 2026-07-17 Kimi K3 · 93.5 · 2026-07-17 Inkling-Small · 89.5 · 2026-08-03 Qwen3.8-Max · 92.6 · 2026-08-03 OpenAI o3 · 87.7 · 2025-04-16
GPT-4 based baseline Claude 3.5 Sonnet Qwen2-72B QwQ-32B-Preview o3 Grok 3 (Think) Claude 3.7 Sonnet Grok 3 Beta DeepSeek R1 OpenAI o3-mini GPT-4.5 Grok 3 DeepSeek-V3-0324 Gemini 2.5 Pro (experimental) Gemini 2.5 Pro Seed1.5-Thinking GPT‑4.1 nano Gemini 2.0 Flash Claude Sonnet 4 Claude Opus 4 Deepseek-R1-0528 Kimi K2 GPT-5 pro GPT-5 Pro DeepSeek-V3.1-Terminus Claude Sonnet 4.5 GPT-5.2 GPT-4 Claude 3 Opus O1 Claude Opus 4.6 Gemini 3 Pro Gemini 3.1 Pro Preview Aristotle-X1 Gemini 3 Deep Think GPT-5.2 Thinking GPT-5.2 Pro Kimi K2.5 Grok 4 GPT-Live-1 Kimi K3 Inkling-Small Qwen3.8-Max OpenAI o3
Timeline
Date Model Score Source
2026-08-03 Qwen3.8-Max 92.6% Alibaba releases Qwen3.8-Max, a 2.4T-parameter MoE model
2026-08-03 Inkling-Small 89.5% Thinking Machines Lab releases Inkling-Small, a 276B open weights multimodal MoE model
2026-07-17 Kimi K3 93.5% Moonshot AI releases open Kimi K3, a 2.8T-parameter MoE model with 1M context
2026-07-17 GPT-Live-1 84.2% OpenAI releases full-duplex voice models and German court holds Google liable for AI Overviews
2026-03-25 o3 87.7% OpenAI announces o3 retirement from ChatGPT and shares benchmark scores
2026-02-25 Grok 4 87.7% xAI releases Grok 4 reasoning model with strong agentic and coding benchmarks
2026-02-05 Claude Opus 4.6 91.3% Anthropic releases Claude Opus 4.6 with 1M context window and agentic improvements
2026-01-27 Kimi K2.5 87.6% Moonshot releases Kimi K2.5 with Agent Swarm technology
2025-12-11 GPT-5.2 Thinking 92.4% OpenAI introduces GPT-5.2 with improved coding, long-context, and reasoning capabilities
2025-12-11 GPT-5.2 Thinking 92.4% OpenAI releases GPT-5.2 Pro and Thinking models for science and math
2025-12-11 GPT-5.2 Pro 93.2% OpenAI releases GPT-5.2 Pro and Thinking models for science and math
2025-12-11 GPT-5.2 Pro 93.2% OpenAI introduces GPT-5.2 with improved coding, long-context, and reasoning capabilities
2025-12-04 Gemini 3 Pro 93.0% Epoch AI launches Frontier Data Centers Hub and analyzes OSWorld benchmark
2025-11-18 Gemini 3 Pro 91.9% Google releases Gemini 3 Pro with state-of-the-art reasoning and multimodal capabilities
2025-11-18 Gemini 3 Deep Think 93.8% Google releases Gemini 3 Pro with state-of-the-art reasoning and multimodal capabilities
2025-11-18 Gemini 3 Pro 91.9% Google releases Gemini 3 Pro with state-of-the-art reasoning and multimodal capabilities
2025-11-18 Gemini 3 Deep Think 93.8% Google releases Gemini 3 Pro with state-of-the-art reasoning and multimodal capabilities
2025-10-24 GPT-5.2 92.4% GPQA-Diamond Benchmark: Scores, Leaderboard & How AI Models Compare
2025-10-24 GPT-4 39.0% GPQA-Diamond Benchmark: Scores, Leaderboard & How AI Models Compare
2025-10-24 Claude 3 Opus 60.0% GPQA-Diamond Benchmark: Scores, Leaderboard & How AI Models Compare
2025-10-24 O1 77.0% GPQA-Diamond Benchmark: Scores, Leaderboard & How AI Models Compare
2025-10-24 Claude Opus 4.6 91.3% GPQA-Diamond Benchmark: Scores, Leaderboard & How AI Models Compare
2025-10-24 Gemini 3 Pro 91.9% GPQA-Diamond Benchmark: Scores, Leaderboard & How AI Models Compare
2025-10-24 Gemini 3.1 Pro Preview 94.1% GPQA-Diamond Benchmark: Scores, Leaderboard & How AI Models Compare
2025-10-24 Aristotle-X1 92.4% GPQA-Diamond Benchmark: Scores, Leaderboard & How AI Models Compare
2025-09-30 Claude Sonnet 4.5 83.4% Anthropic releases Claude Sonnet 4.5 with improved coding and agent capabilities
2025-09-22 DeepSeek-V3.1-Terminus 74.24% DeepSeek releases DeepSeek-V3.1-Terminus with language and agent fixes
2025-08-07 GPT-5 pro 88.4% OpenAI introduces GPT-5 with unified routing and expert-level reasoning
2025-08-07 GPT-5 Pro 88.4% OpenAI launches GPT-5 with adaptive reasoning and unified architecture
2025-08-07 GPT-5 pro 89.4% OpenAI launches unified GPT-5 with real-time routing and free access
2025-07-28 Kimi K2 75.1% Kimi K2: open MoE model with 1T parameters and MuonClip optimizer
2025-07-28 Kimi K2 75.1% Moonshot releases Kimi K2, an open agentic MoE model with 32B activated parameters
2025-05-30 Deepseek-R1-0528 81.0% DeepSeek updates R1 model with improved reasoning and benchmark scores
2025-05-22 Claude Sonnet 4 70.0% Anthropic introduces Claude Opus 4 and Sonnet 4 with extended thinking and tool use
2025-05-22 Claude Opus 4 74.9% Anthropic releases Claude Opus 4 and Sonnet 4 with ASL-3 safety measures
2025-05-22 Claude Opus 4 74.9% Anthropic introduces Claude Opus 4 and Sonnet 4 with extended thinking and tool use
2025-04-16 OpenAI o3 87.7% OpenAI
2025-04-15 Gemini 2.0 Flash 60.1% Google releases Gemini 2.0 Flash with enhanced quality and twice the speed of Gemini 1.5 Pro
2025-04-14 GPT‑4.1 nano 50.3% OpenAI launches GPT-4.1, GPT-4.1 mini, and GPT-4.1 nano in the API
2025-04-10 Seed1.5-Thinking 77.3% Seed1.5-Thinking uses reinforcement learning to improve reasoning
2025-04-05 Gemini 2.5 Pro 84.0% Google expands access to Gemini 2.5 Pro amid strong benchmark results
2025-03-26 Gemini 2.5 Pro 84.0% Google releases experimental Gemini 2.5 Pro reasoning model with 1M token context
2025-03-25 Gemini 2.5 Pro (experimental) 84.0% Google DeepMind releases Gemini 2.5 Pro experimental model
2025-03-25 Gemini 2.5 Pro 84.0% Google DeepMind
2025-03-24 DeepSeek-V3-0324 68.4% DeepSeek-V3-0324 improves reasoning, coding, and Chinese writing over DeepSeek-V3
2025-03-11 Grok 3 84.6% xAI releases Grok 3 with deep reasoning and million-token context
2025-03-05 GPT-4.5 71.4% OpenAI releases GPT-4.5, its largest model, for ChatGPT Plus
2025-02-26 Claude 3.7 Sonnet 84.8% Benchmark comparison of Claude 3.7 Sonnet, o3-mini, R1, and Grok 3 Beta
2025-02-26 Grok 3 Beta 84.6% Benchmark comparison of Claude 3.7 Sonnet, o3-mini, R1, and Grok 3 Beta
2025-02-26 DeepSeek R1 71.5% Benchmark comparison of Claude 3.7 Sonnet, o3-mini, R1, and Grok 3 Beta
2025-02-26 OpenAI o3-mini 78.0% Benchmark comparison of Claude 3.7 Sonnet, o3-mini, R1, and Grok 3 Beta
2025-02-26 Claude 3.5 Sonnet 65.0% Benchmark comparison of Claude 3.7 Sonnet, o3-mini, R1, and Grok 3 Beta
2025-02-24 Claude 3.7 Sonnet 84.8% Anthropic releases Claude 3.7 Sonnet with dynamic reasoning control
2025-02-19 Grok 3 (Think) 84.6% xAI releases Grok 3 Beta and DeepSearch agent
2024-12-20 o3 87.7% OpenAI releases o3 model with high performance in reasoning and coding
2024-11-28 QwQ-32B-Preview 65.2% Qwen releases QwQ-32B-Preview, an experimental reasoning model
2024-07-15 Qwen2-72B 37.9% Qwen Team releases Qwen2 series with 72B dense and MoE models
2024-07-15 Qwen2-72B 37.9% Qwen2 Technical Report introduces dense and MoE models up to 72B
2024-06-20 Claude 3.5 Sonnet 59.4% Anthropic releases Claude 3.5 Sonnet addendum with improved coding and vision benchmarks
2024-06-20 Claude 3.5 Sonnet 59.4% Anthropic
2023-11-20 GPT-4 based baseline 39.0% GPQA: A Graduate-Level Google-Proof Q&A Benchmark