MMLU
MMLU is a 57-task test of a text model's multitask accuracy.
MMLU is a collection of multiple-choice tasks. Its value is breadth: the questions span 57 academic and professional subjects rather than concentrating on one topic.
The evaluation path is simple. A system receives a question and answer options, returns one option, and is checked against a reference answer. Scores can then be summarized by subject category and as an overall average.
Interpret the number narrowly. The original authors found lopsided performance and reported that models frequently did not know when they were wrong. MMLU-CF identifies contamination as a concern for public multiple-choice benchmarks. MMLU-Pro expands the answer choices from four to ten. Record the benchmark version, prompt and scoring setup with each result.
A single subject cannot show how broadly a model answers knowledge questions.
Present questions, collect option choices, and aggregate accuracy.
- 1 · presentGive the model a question and its answer options.
- 2 · chooseRecord one selected option for the question.
- 3 · scoreCompare the selection with the reference answer and aggregate accuracy.
The original authors report that models can have lopsided performance and frequently do not know when they are wrong.
| Who | What they ask | What it works with |
|---|---|---|
| Model evaluator | “How accurate is this model across many knowledge subjects?” | Subject-level and average accuracy |
| Researcher | “Which broad areas are weaker?” | Category and task results |
| Model developer | “Does a training change alter the same fixed evaluation?” | Results under the same prompt and scoring setup |
- Tests multiple-choice performance across 57 tasks.
- Includes subjects such as mathematics, history, computer science and law.
- Produces category results and an overall average in the reference repository.
- Supports analysis of performance breadth and depth across tasks.
- MMLU does not show that a model knows when it is wrong.
- High average accuracy can coexist with lopsided task performance.
- The original test still required substantial improvement to reach expert-level accuracy on every task studied.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperMeasuring Massive Multitask Language Understanding, Hendrycks and colleagues · read 28 Sept 2026
- repoMeasuring Massive Multitask Language Understanding repository, MMLU authors · read 28 Sept 2026
- paperMMLU-CF A Contamination-free Multi-task Language Understanding Benchmark, Zhao and colleagues · read 28 Sept 2026
- paperMMLU-Pro, Wang and colleagues · read 28 Sept 2026
- repoLanguage Model Evaluation Harness, EleutherAI · read 28 Sept 2026