ARC-AGI
ARC-AGI presents demonstration grid pairs to a system, then asks it to construct outputs for new test grids.
ARC-AGI presents small colored-grid puzzles. Demonstration pairs show how inputs become outputs. A system must infer a rule that fits those pairs and use it to construct the output for a new test grid.
The highlighted step is rule inference. A submitted output counts as correct only when every cell matches the expected answer.
Name the version when reporting a result. ARC-AGI-2 preserves the input-output pair format and contains a newly curated and expanded task set. ARC-AGI-1 and ARC-AGI-2 scores refer to different task sets.
The benchmark targets skill acquisition from sparse demonstrations.
Compare example pairs, infer one rule, then construct the test output grid.
- 1 · observeInspect the input-output demonstration pairs.
- 2 · inferForm a transformation rule consistent with those pairs.
- 3 · applyConstruct the output grid for the held-out test input.
The original repository says each task has demonstration examples followed by one or more test inputs.
| Who | What they ask | What it works with |
|---|---|---|
| Reasoning researcher | “Can a system acquire a new grid transformation from few examples?” | Exact task success under the stated attempt policy |
| Benchmark operator | “Are evaluation answers kept separate from public training tasks?” | Dataset split and submission procedure |
| Result reader | “Which ARC version produced this score?” | ARC-AGI version and evaluation rules |
- Uses grids of integer symbols that are visualized as colors.
- Shows demonstrations before test inputs within each task.
- Checks whether the submitted output grid exactly matches the expected grid.
- Separates public training tasks from evaluation tasks.
- The ARC paper says skill at one task falls short of measuring intelligence.
- The paper says skill is strongly affected by prior knowledge and experience.
- Scores from ARC-AGI-1 and ARC-AGI-2 refer to different task sets.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperOn the Measure of Intelligence, François Chollet · read 28 Sept 2026
- repoARC-AGI repository, François Chollet · read 28 Sept 2026
- officialARC-AGI-1, ARC Prize Foundation · read 28 Sept 2026
- officialARC-AGI-2, ARC Prize Foundation · read 28 Sept 2026
- paperARC-AGI-2 A New Challenge for Frontier AI Reasoning Systems, Chollet and colleagues · read 28 Sept 2026