Benchmark · general

MMLU

saturated 15 results 12 models

MMLU (Massive Multitask Language Understanding) tests a model's breadth of knowledge across 57 subjects—from history and law to math and medicine—using multiple-choice questions. Performance is reported as % accuracy: the share of questions answered correctly.

Read more
Example
A typical item gives a question and four labeled answer choices (A–D), and the model must pick the correct one—for instance, a college-level medicine question or an elementary math problem.
Scoring
The metric is accuracy: the percentage of multiple-choice questions for which the model selects the correct option, averaged across all 57 subjects.
Verification
Grading is automatic and objective—the model's chosen letter is exact-matched against the known correct answer, so no human judgment is needed.
Why it matters
MMLU became a standard yardstick for general knowledge and reasoning breadth, but top models now score so high that it is largely saturated and less useful for separating the best systems.
Worked example
Task
The following is a multiple-choice question about high school chemistry. Which of the following is a strong acid? A) HF B) HCl C) CH3COOH D) H2CO3
Solution
HCl ionizes completely in water, so it is a strong acid; HF, CH3COOH, and H2CO3 ionize only partially (weak acids). Final answer: B) HCl.
Walkthrough
A strong acid fully dissociates into H+ and its conjugate base in aqueous solution, which HCl does and the others do not. MMLU grades it by exact-match accuracy: the predicted option letter must equal the gold label (B).
0 24 48 72 96 2023-03-14 2024-03-29 2025-04-15 Gemini Ultra · 90 · 2023-12-06 Gemini Ultra · 90 · 2023-12-06 Qwen2-72B · 84.2 · 2024-07-15 Qwen2-72B · 84.2 · 2024-07-15 Mistral Large 2 · 84 · 2024-07-24 Mistral Large 2 · 84 · 2024-07-25 Falcon3-10B-Base · 73.1 · 2024-12-17 Falcon3-7B-Base · 67.4 · 2024-12-17 Claude 3.7 Sonnet · 86.1 · 2025-02-24 GPT‑4.1 nano · 80.1 · 2025-04-14 Gemini 2.0 Flash · 83.4 · 2025-04-15 GPT-3.5 · 70 · 2023-03-14 GPT-4 · 86.4 · 2023-03-14 Claude 2 · 78.5 · 2023-07-11 Claude 3 Opus · 86.8 · 2024-03-04
Gemini Ultra Qwen2-72B Mistral Large 2 Falcon3-10B-Base Falcon3-7B-Base Claude 3.7 Sonnet GPT‑4.1 nano Gemini 2.0 Flash GPT-3.5 GPT-4 Claude 2 Claude 3 Opus
Timeline
Date Model Score Source
2025-04-15 Gemini 2.0 Flash 83.4% Google releases Gemini 2.0 Flash with enhanced quality and twice the speed of Gemini 1.5 Pro
2025-04-14 GPT‑4.1 nano 80.1% OpenAI launches GPT-4.1, GPT-4.1 mini, and GPT-4.1 nano in the API
2025-02-24 Claude 3.7 Sonnet 86.1% Anthropic releases Claude 3.7 Sonnet with dynamic reasoning control
2024-12-17 Falcon3-10B-Base 73.1% Falcon3 family releases five open models with improved science, math, and code capabilities
2024-12-17 Falcon3-7B-Base 67.4% Falcon3 family releases five open models with improved science, math, and code capabilities
2024-07-25 Mistral Large 2 84.0% Mistral releases Mistral Large 2 with 128K context and improved multilingual support
2024-07-24 Mistral Large 2 84.0% Mistral releases Large 2 with 123B parameters for cost-efficient single-node inference
2024-07-15 Qwen2-72B 84.2% Qwen Team releases Qwen2 series with 72B dense and MoE models
2024-07-15 Qwen2-72B 84.2% Qwen2 Technical Report introduces dense and MoE models up to 72B
2024-03-04 Claude 3 Opus 86.8% Anthropic
2023-12-06 Gemini Ultra 90.0% Google introduces Gemini 1.0, its largest multimodal AI model
2023-12-06 Gemini Ultra 90.0% Google
2023-07-11 Claude 2 78.5% Anthropic
2023-03-14 GPT-3.5 70.0% OpenAI
2023-03-14 GPT-4 86.4% OpenAI