Concepts

GPQA

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

GPQA is a 448-question multiple-choice benchmark written by experts in biology, physics and chemistry.

1 · What it is

GPQA is a set of 448 expert-written multiple-choice questions in biology, physics and chemistry. The paper reports both expert and skilled non-expert baselines, making the gap between subject expertise and general research skill visible.

To run it, present a question and its options, record one answer, and compare it with the validated answer key. The authors’ repository also exposes a seed for shuffling answer choices, so an evaluation report should state the split and configuration.

The result is accuracy on this collection. The authors connect its difficulty for skilled non-experts and frontier systems to realistic scalable-oversight experiments.

2 · Why it exists

Some science questions are hard to supervise without matching expertise.

DifficultyGPQA contains difficult graduate-level questions in three scientific fields.
OversightThe authors designed GPQA to support experiments on supervising answers to hard questions.
BaselineThe paper reports expert, skilled non-expert and AI-system baselines.
3 · How it works

Present the question, choose an option, then score against the expert key.

GPQA reports accuracy on 448 expert-written science questions.
  1. 1 · questionPresent an expert-written biology, physics or chemistry question.
  2. 2 · chooseRecord one answer from the multiple-choice options.
  3. 3 · scoreCompare that answer with the validated key and compute accuracy.

The authors call the questions Google-proof because skilled non-experts struggled despite unrestricted web access.

4 · Where it's used
WhoWhat they askWhat it works with
Evaluation researcher“Can a system answer difficult graduate science questions?”GPQA split and scoring protocol
Oversight researcher“Can a less expert evaluator supervise a stronger answerer?”Expert and non-expert baselines
Benchmark operator“Which subset and answer-order policy were used?”Dataset configuration and seed
5 · What it solves, and what it doesn't
solves
  • Supplies 448 expert-written multiple-choice science questions.
  • Covers biology, physics and chemistry.
  • Reports a 65 percent expert baseline in the original study.
  • Provides baseline code and analysis in the authors' repository.
doesn't solve
  • Skilled non-expert validators reached 34 percent in the original study despite unrestricted web access.
  • Experts reached 65 percent in the original study, so the expert baseline was not perfect.
  • The paper's strongest GPT-4-based baseline reached 39 percent at publication time.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperGPQA A Graduate-Level Google-Proof Q&A Benchmark, Rein and colleagues · read 28 Sept 2026
  2. repoGPQA repository, GPQA authors · read 28 Sept 2026
  3. repoSimple Evals, OpenAI · read 28 Sept 2026
  4. paperOn scalable oversight with weak LLMs judging strong LLMs, Kenton and colleagues · read 28 Sept 2026
  5. repoLanguage Model Evaluation Harness GPQA task, EleutherAI · read 28 Sept 2026