Concepts

ARC-AGI

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

ARC-AGI presents demonstration grid pairs to a system, then asks it to construct outputs for new test grids.

1 · What it is

ARC-AGI presents small colored-grid puzzles. Demonstration pairs show how inputs become outputs. A system must infer a rule that fits those pairs and use it to construct the output for a new test grid.

The highlighted step is rule inference. A submitted output counts as correct only when every cell matches the expected answer.

Name the version when reporting a result. ARC-AGI-2 preserves the input-output pair format and contains a newly curated and expanded task set. ARC-AGI-1 and ARC-AGI-2 scores refer to different task sets.

2 · Why it exists

The benchmark targets skill acquisition from sparse demonstrations.

InferEach task requires finding a transformation from example grid pairs.
ApplyThe inferred transformation must produce the output grid for a new input.
SeparatePublic training tasks and evaluation tasks are distinct sets.
3 · How it works

Compare example pairs, infer one rule, then construct the test output grid.

ARC-AGI evaluates the exact output grid produced for a new input after a few demonstrations.
  1. 1 · observeInspect the input-output demonstration pairs.
  2. 2 · inferForm a transformation rule consistent with those pairs.
  3. 3 · applyConstruct the output grid for the held-out test input.

The original repository says each task has demonstration examples followed by one or more test inputs.

4 · Where it's used
WhoWhat they askWhat it works with
Reasoning researcher“Can a system acquire a new grid transformation from few examples?”Exact task success under the stated attempt policy
Benchmark operator“Are evaluation answers kept separate from public training tasks?”Dataset split and submission procedure
Result reader“Which ARC version produced this score?”ARC-AGI version and evaluation rules
5 · What it solves, and what it doesn't
solves
  • Uses grids of integer symbols that are visualized as colors.
  • Shows demonstrations before test inputs within each task.
  • Checks whether the submitted output grid exactly matches the expected grid.
  • Separates public training tasks from evaluation tasks.
doesn't solve
  • The ARC paper says skill at one task falls short of measuring intelligence.
  • The paper says skill is strongly affected by prior knowledge and experience.
  • Scores from ARC-AGI-1 and ARC-AGI-2 refer to different task sets.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperOn the Measure of Intelligence, François Chollet · read 28 Sept 2026
  2. repoARC-AGI repository, François Chollet · read 28 Sept 2026
  3. officialARC-AGI-1, ARC Prize Foundation · read 28 Sept 2026
  4. officialARC-AGI-2, ARC Prize Foundation · read 28 Sept 2026
  5. paperARC-AGI-2 A New Challenge for Frontier AI Reasoning Systems, Chollet and colleagues · read 28 Sept 2026