GPQA
GPQA is a 448-question multiple-choice benchmark written by experts in biology, physics and chemistry.
GPQA is a set of 448 expert-written multiple-choice questions in biology, physics and chemistry. The paper reports both expert and skilled non-expert baselines, making the gap between subject expertise and general research skill visible.
To run it, present a question and its options, record one answer, and compare it with the validated answer key. The authors’ repository also exposes a seed for shuffling answer choices, so an evaluation report should state the split and configuration.
The result is accuracy on this collection. The authors connect its difficulty for skilled non-experts and frontier systems to realistic scalable-oversight experiments.
Some science questions are hard to supervise without matching expertise.
Present the question, choose an option, then score against the expert key.
- 1 · questionPresent an expert-written biology, physics or chemistry question.
- 2 · chooseRecord one answer from the multiple-choice options.
- 3 · scoreCompare that answer with the validated key and compute accuracy.
The authors call the questions Google-proof because skilled non-experts struggled despite unrestricted web access.
| Who | What they ask | What it works with |
|---|---|---|
| Evaluation researcher | “Can a system answer difficult graduate science questions?” | GPQA split and scoring protocol |
| Oversight researcher | “Can a less expert evaluator supervise a stronger answerer?” | Expert and non-expert baselines |
| Benchmark operator | “Which subset and answer-order policy were used?” | Dataset configuration and seed |
- Supplies 448 expert-written multiple-choice science questions.
- Covers biology, physics and chemistry.
- Reports a 65 percent expert baseline in the original study.
- Provides baseline code and analysis in the authors' repository.
- Skilled non-expert validators reached 34 percent in the original study despite unrestricted web access.
- Experts reached 65 percent in the original study, so the expert baseline was not perfect.
- The paper's strongest GPT-4-based baseline reached 39 percent at publication time.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperGPQA A Graduate-Level Google-Proof Q&A Benchmark, Rein and colleagues · read 28 Sept 2026
- repoGPQA repository, GPQA authors · read 28 Sept 2026
- repoSimple Evals, OpenAI · read 28 Sept 2026
- paperOn scalable oversight with weak LLMs judging strong LLMs, Kenton and colleagues · read 28 Sept 2026
- repoLanguage Model Evaluation Harness GPQA task, EleutherAI · read 28 Sept 2026