Benchmark · reasoning

ARC-AGI 1

28 results 26 models

ARC-AGI 1 measures abstract reasoning and generalization: each task shows a few input→output grid examples, and the solver must infer the hidden transformation rule and produce the correct output grid for a new input. The metric is the percentage of tasks solved with an exactly matching output grid.

Read more
Example
A task might present two or three small colored grids where each input turns into its output by one consistent rule — for instance filling every enclosed region or mirroring a shape — and you must apply that inferred rule to a new, unseen grid.
Scoring
A task counts as solved only when the produced grid matches the reference answer cell-for-cell; the final score is the fraction (percentage) of tasks solved across the evaluation set.
Verification
Checking is automatic and exact: the submitted output grid is compared cell-by-cell with the ground-truth grid, with no partial credit and no human voting.
Why it matters
The puzzles are built to be easy for people but hard for models that lean on memorized patterns, so the benchmark probes real abstraction and few-shot generalization rather than recalled knowledge — making it a closely watched measure of progress toward more general AI.
Worked example
Task
ARC-AGI 1 shows a few input→output grid demonstrations, then a held-out test input; you must infer the transformation and draw the exact output grid (cells are colors 0-9, 0=black). Here the demos reveal a self-fractal rule: each 3×3 grid expands to 9×9, stamping a copy of the whole grid wherever a cell is colored and a blank 3×3 block wherever it is 0. Test input (3×3): 0 8 0 / 8 8 8 / 0 8 0 (a plus shape in color 8).
Solution
0 0 0 0 8 0 0 0 0 0 0 0 8 8 8 0 0 0 0 0 0 0 8 0 0 0 0 0 8 0 0 8 0 0 8 0 8 8 8 8 8 8 8 8 8 0 8 0 0 8 0 0 8 0 0 0 0 0 8 0 0 0 0 0 0 0 8 8 8 0 0 0 0 0 0 0 8 0 0 0 0
Walkthrough
The color-8 cells sit at the four edge-midpoints and the center of the plus; stamping the plus into exactly those five 3×3 blocks (and leaving the four corners blank) builds the 9×9 fractal. ARC-AGI scores an exact cell-for-cell grid match — a single wrong cell fails the task, with up to 2 attempts allowed.
0 24.5 49 73.5 98 2024-06-17 2025-06-21 2026-06-25 GPT-4o · 50 · 2024-06-17 MindsAI · 55.5 · 2024-12-06 Sonnet 3.5.1 · 53.6 · 2024-12-06 MIT & Cornell team · 47.5 · 2024-12-18 Claude Sonnet 3.5 · 53.6 · 2024-12-18 o3 tuned low · 76 · 2024-12-20 o3 · 75.7 · 2024-12-20 o3 · 75.7 · 2024-12-20 o3 · 75 · 2026-04-16 o3-preview (low) · 75.7 · 2025-04-20 o3-preview (high) · 87.5 · 2025-04-20 o3-preview-low · 76 · 2025-04-22 o4-mini-medium · 41 · 2025-04-22 o4-mini-low · 21 · 2025-04-22 o3-low · 41 · 2025-04-22 o3-medium · 53 · 2025-04-22 o3-preview-high · 88 · 2025-04-22 HRM · 32 · 2025-08-15 Grok-4 · 79.6 · 2025-09-16 TRM · 45 · 2025-10-06 Tiny Recursive Model (TRM) · 45 · 2025-12-05 CompressARC · 20 · 2025-12-05 GPT-5.2 Thinking · 86.2 · 2025-12-11 GPT-5.2 Pro · 90.5 · 2025-12-11 Opus 4.6 · 93 · 2026-03-09 iLLaDA-Base · 14.9 · 2026-06-25 OpenAI o3 (low compute) · 75.7 · 2024-12-20 OpenAI o3 (high compute) · 87.5 · 2024-12-20
GPT-4o MindsAI Sonnet 3.5.1 MIT & Cornell team Claude Sonnet 3.5 o3 tuned low o3 o3-preview (low) o3-preview (high) o3-preview-low o4-mini-medium o4-mini-low o3-low o3-medium o3-preview-high HRM Grok-4 TRM Tiny Recursive Model (TRM) CompressARC GPT-5.2 Thinking GPT-5.2 Pro Opus 4.6 iLLaDA-Base OpenAI o3 (low compute) OpenAI o3 (high compute)
Timeline
Date Model Score Source
2026-06-25 iLLaDA-Base 14.9pts iLLaDA: An 8B Masked Diffusion Language Model with Fully Bidirectional Attention
2026-04-16 o3 75.0% GPT-5.2 achieves ~53% on ARC-AGI-2 benchmark in December 2025
2026-03-09 Opus 4.6 93.0% Survey finds 2-3x performance drops on ARC-AGI benchmarks despite cost falling 390x
2025-12-11 GPT-5.2 Thinking 86.2% OpenAI introduces GPT-5.2 with improved coding, long-context, and reasoning capabilities
2025-12-11 GPT-5.2 Pro 90.5% OpenAI releases GPT-5.2, outperforming Gemini 3 in coding and reasoning benchmarks
2025-12-05 Tiny Recursive Model (TRM) 45.0% ARC Prize 2025 results highlight refinement loops as key to AGI progress
2025-12-05 CompressARC 20.0% ARC Prize 2025 results highlight refinement loops as key to AGI progress
2025-10-06 TRM 45.0% Tiny Recursive Model (TRM) achieves higher ARC-AGI accuracy than LLMs with 7M parameters
2025-09-16 Grok-4 79.6% ARC-AGI v2 SoTA set by evolving English instructions with Grok-4
2025-08-15 HRM 32.0% Analysis finds HRM's ARC-AGI performance driven by outer loop refinement and memorization
2025-04-22 o3-preview-low 76.0% ARC Prize Foundation evaluates OpenAI o3 and o4-mini on ARC-AGI benchmarks
2025-04-22 o4-mini-medium 41.0% ARC Prize Foundation evaluates OpenAI o3 and o4-mini on ARC-AGI benchmarks
2025-04-22 o4-mini-low 21.0% ARC Prize Foundation evaluates OpenAI o3 and o4-mini on ARC-AGI benchmarks
2025-04-22 o3-low 41.0% ARC Prize Foundation evaluates OpenAI o3 and o4-mini on ARC-AGI benchmarks
2025-04-22 o3-medium 53.0% ARC Prize Foundation evaluates OpenAI o3 and o4-mini on ARC-AGI benchmarks
2025-04-22 o3-preview-high 88.0% ARC Prize Foundation evaluates OpenAI o3 and o4-mini on ARC-AGI benchmarks
2025-04-20 o3-preview (low) 75.7% OpenAI's o3 scores lower on FrontierMath than initially claimed
2025-04-20 o3-preview (high) 87.5% OpenAI's o3 scores lower on FrontierMath than initially claimed
2024-12-20 o3 tuned low 76.0% OpenAI previews o3 reasoning model with ARC AGI and Frontier Math breakthroughs
2024-12-20 o3 75.7% OpenAI o3 achieves breakthrough scores on ARC-AGI-Pub
2024-12-20 o3 75.7% OpenAI releases o3 model with high performance in reasoning and coding
2024-12-20 OpenAI o3 (low compute) 75.7% ARC Prize
2024-12-20 OpenAI o3 (high compute) 87.5% ARC Prize
2024-12-18 MIT & Cornell team 47.5% Jeremy Berman and MIT/Cornell team achieve new state-of-the-art on ARC-AGI-Pub
2024-12-18 Claude Sonnet 3.5 53.6% Jeremy Berman and MIT/Cornell team achieve new state-of-the-art on ARC-AGI-Pub
2024-12-06 MindsAI 55.5% ARC Prize 2024 winners published; SOTA on ARC-AGI-1 rises to 55.5%
2024-12-06 Sonnet 3.5.1 53.6% Jeremy Berman achieves 53.6% on ARC-AGI-Pub using Sonnet 3.5 with Evolutionary Test-time Compute
2024-06-17 GPT-4o 50.0% GPT-4o achieves 50% SoTA on ARC-AGI via massive sampling