Benchmark · reasoning

Humanity's Last Exam

32 results 24 models

Humanity's Last Exam (HLE) is a benchmark of frontier, expert-level questions spanning many academic domains (mathematics, sciences, humanities, and more), deliberately built to be extremely hard for AI. It reports performance as percent accuracy.

Read more
Example
Each item is a closed-ended expert question with a single correct answer — for instance, an advanced mathematics problem or a graduate-level science question (some with an accompanying image), posed as multiple choice or an exact short answer.
Scoring
The metric is accuracy: the percentage of questions the model answers correctly. Some versions also report a calibration (confidence) measure alongside accuracy.
Verification
Each question has a known correct answer, and a model's response is graded automatically against that reference (the correct option for multiple choice, a matching short answer) — objective automated grading, not human votes.
Why it matters
Standard benchmarks have become saturated (top models score near the ceiling), so HLE aims to be a hard, discriminating test at the frontier of human expert knowledge — a meaningful gauge of how far AI is from expert-level reasoning.
Worked example
Task
Hummingbirds (order Apodiformes) uniquely possess a bilaterally paired, oval sesamoid bone embedded in the caudolateral part of the expanded, cruciate aponeurosis of insertion of the m. depressor caudae. How many paired tendons does this sesamoid bone support? Answer with a single integer.
Solution
4
Walkthrough
The oval sesamoid forms within the cruciate aponeurosis of the tail-depressor muscle and supports four paired tendons of that insertion — a specialized ossification unique to hummingbird tail anatomy. HLE grades the response by exact match of the submitted integer against the reference answer (4) and separately elicits a stated confidence for calibration scoring.
0 16 32 48 64 2025-01-23 2025-11-02 2026-08-13 the model powering deep research · 26.6 · 2025-02-02 Gemini 2.5 Pro (experimental) · 18.8 · 2025-03-25 Gemini 2.5 Pro · 18.8 · 2025-03-25 Gemini 2.5 Pro · 18.8 · 2025-03-25 Gemini 2.5 Pro · 18.8 · 2025-03-26 Gemini 2.5 Pro · 18.8 · 2025-04-05 Grok 4 Heavy · 50.7 · 2025-07-09 GPT-5 Pro · 42 · 2025-08-07 Tongyi DeepResearch · 32.9 · 2025-09-16 DeepSeek-V3.2-Exp · 56.5 · 2025-09-29 Kimi K2 Thinking · 44.9 · 2025-11-07 MiroThinker v1.0 (72B variant) · 37.7 · 2025-11-14 Gemini 3 Pro · 37.5 · 2025-11-18 Gemini 3 Pro · 37.5 · 2025-11-18 Gemini 3 Deep Think · 41 · 2025-11-18 Gemini 3 Deep Think · 41 · 2025-11-18 Gemini 3 Deep Think · 48.4 · 2026-02-12 Kimi K2.5 · 51.8 · 2026-01-29 Claude Opus 4.6 · 53 · 2026-02-05 Claude Opus 4.6 · 40 · 2026-02-05 Claude Opus 4.6 · 53 · 2026-03-06 Claude Opus 4.7 · 46.9 · 2026-04-25 GPT-5.5 · 41.4 · 2026-04-25 Claude Fable 5 · 53 · 2026-06-09 Agents-A1 · 47.6 · 2026-06-30 Claude Opus 4.8 · 46 · 2026-07-01 PoTRE · 49.9 · 2026-07-23 Inkling-Small · 31.6 · 2026-08-03 DeepSeek-V4-Pro · 60 · 2026-08-13 DeepSeek R1 (text-only) · 9.4 · 2025-01-23 Grok 4 · 25.4 · 2025-07-09 GPT-5 · 24.8 · 2025-08-07
the model powering deep research Gemini 2.5 Pro (experimental) Gemini 2.5 Pro Grok 4 Heavy GPT-5 Pro Tongyi DeepResearch DeepSeek-V3.2-Exp Kimi K2 Thinking MiroThinker v1.0 (72B variant) Gemini 3 Pro Gemini 3 Deep Think Kimi K2.5 Claude Opus 4.6 Claude Opus 4.7 GPT-5.5 Claude Fable 5 Agents-A1 Claude Opus 4.8 PoTRE Inkling-Small DeepSeek-V4-Pro DeepSeek R1 (text-only) Grok 4 GPT-5
Timeline
Date Model Score Source
2026-08-13 DeepSeek-V4-Pro 60.0% DeepSeek-V4-Pro GA release enhances agent capabilities and adds Responses API support
2026-08-03 Inkling-Small 31.6% Thinking Machines Lab releases Inkling-Small, a 276B open weights multimodal MoE model
2026-07-23 PoTRE 49.92% PoTRE introduces heterogeneous multi-agent framework for test-time reasoning
2026-07-01 Claude Opus 4.8 46.0% Anthropic releases Claude Opus 4.8 with adaptive reasoning and improved honesty
2026-06-30 Agents-A1 47.6pts Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent
2026-06-09 Claude Fable 5 53.0% Claude Fable 5 launches at #1 on the Artificial Analysis Intelligence Index
2026-04-25 Claude Opus 4.7 46.9% OpenAI releases GPT-5.5, a ground-up retrained model with native omnimodal processing
2026-04-25 GPT-5.5 41.4% OpenAI releases GPT-5.5, a ground-up retrained model with native omnimodal processing
2026-03-06 Claude Opus 4.6 53.0% Anthropic releases Claude Opus 4.6 system card detailing capabilities and safety assessments
2026-02-12 Gemini 3 Deep Think 48.4% Google releases upgraded Gemini 3 Deep Think with new benchmark scores
2026-02-05 Claude Opus 4.6 53.0% Anthropic publishes Claude Opus 4.6 system card detailing capabilities and safety evaluations
2026-02-05 Claude Opus 4.6 40.0% Anthropic releases Claude Opus 4.6 with 1M context window and agentic improvements
2026-01-29 Kimi K2.5 51.8% Moonshot releases Kimi K2.5 multimodal agentic model
2025-11-18 Gemini 3 Pro 37.5% Google releases Gemini 3 Pro with state-of-the-art reasoning and multimodal capabilities
2025-11-18 Gemini 3 Deep Think 41.0% Google releases Gemini 3 Pro with state-of-the-art reasoning and multimodal capabilities
2025-11-18 Gemini 3 Pro 37.5% Google releases Gemini 3 Pro with state-of-the-art reasoning and multimodal capabilities
2025-11-18 Gemini 3 Deep Think 41.0% Google releases Gemini 3 Pro with state-of-the-art reasoning and multimodal capabilities
2025-11-14 MiroThinker v1.0 (72B variant) 37.7% MiroThinker v1.0 introduces interactive scaling for open-source research agents
2025-11-07 Kimi K2 Thinking 44.9% Moonshot AI releases Kimi K2 Thinking, an open-source 1T parameter reasoning model
2025-09-29 DeepSeek-V3.2-Exp 56.53% DeepSeek releases DeepSeek-V3.2-Exp with Sparse Attention for long-context efficiency
2025-09-16 Tongyi DeepResearch 32.9% Tongyi releases DeepResearch, an open-source web agent matching proprietary performance
2025-08-07 GPT-5 Pro 42.0% OpenAI launches unified GPT-5 with real-time routing and free access
2025-08-07 GPT-5 24.8% OpenAI
2025-07-09 Grok 4 Heavy 50.7% xAI releases Grok 4 with scaled reinforcement learning and native tool use
2025-07-09 Grok 4 25.4% xAI
2025-04-05 Gemini 2.5 Pro 18.8% Google expands access to Gemini 2.5 Pro amid strong benchmark results
2025-03-26 Gemini 2.5 Pro 18.8% Google releases experimental Gemini 2.5 Pro reasoning model with 1M token context
2025-03-25 Gemini 2.5 Pro (experimental) 18.8% Google DeepMind releases Gemini 2.5 Pro experimental model
2025-03-25 Gemini 2.5 Pro 18.8% Google introduces Gemini 2.5 Pro thinking model
2025-03-25 Gemini 2.5 Pro 18.8% Google introduces Gemini 2.5 Pro experimental model
2025-02-02 the model powering deep research 26.6% OpenAI launches deep research in ChatGPT as an agentic capability
2025-01-23 DeepSeek R1 (text-only) 9.4% Humanity's Last Exam