Concepts

MMLU

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

MMLU is a 57-task test of a text model's multitask accuracy.

1 · What it is

MMLU is a collection of multiple-choice tasks. Its value is breadth: the questions span 57 academic and professional subjects rather than concentrating on one topic.

The evaluation path is simple. A system receives a question and answer options, returns one option, and is checked against a reference answer. Scores can then be summarized by subject category and as an overall average.

Interpret the number narrowly. The original authors found lopsided performance and reported that models frequently did not know when they were wrong. MMLU-CF identifies contamination as a concern for public multiple-choice benchmarks. MMLU-Pro expands the answer choices from four to ten. Record the benchmark version, prompt and scoring setup with each result.

2 · Why it exists

A single subject cannot show how broadly a model answers knowledge questions.

BreadthMMLU covers 57 tasks from several academic and professional fields.
ComparisonThe repository reports category scores and an overall average.
DiagnosisThe original paper uses the test to identify shortcomings across tasks.
3 · How it works

Present questions, collect option choices, and aggregate accuracy.

MMLU converts answers across 57 tasks into accuracy summaries; the score describes this test, not every kind of intelligence.
  1. 1 · presentGive the model a question and its answer options.
  2. 2 · chooseRecord one selected option for the question.
  3. 3 · scoreCompare the selection with the reference answer and aggregate accuracy.

The original authors report that models can have lopsided performance and frequently do not know when they are wrong.

4 · Where it's used
WhoWhat they askWhat it works with
Model evaluator“How accurate is this model across many knowledge subjects?”Subject-level and average accuracy
Researcher“Which broad areas are weaker?”Category and task results
Model developer“Does a training change alter the same fixed evaluation?”Results under the same prompt and scoring setup
5 · What it solves, and what it doesn't
solves
  • Tests multiple-choice performance across 57 tasks.
  • Includes subjects such as mathematics, history, computer science and law.
  • Produces category results and an overall average in the reference repository.
  • Supports analysis of performance breadth and depth across tasks.
doesn't solve
  • MMLU does not show that a model knows when it is wrong.
  • High average accuracy can coexist with lopsided task performance.
  • The original test still required substantial improvement to reach expert-level accuracy on every task studied.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperMeasuring Massive Multitask Language Understanding, Hendrycks and colleagues · read 28 Sept 2026
  2. repoMeasuring Massive Multitask Language Understanding repository, MMLU authors · read 28 Sept 2026
  3. paperMMLU-CF A Contamination-free Multi-task Language Understanding Benchmark, Zhao and colleagues · read 28 Sept 2026
  4. paperMMLU-Pro, Wang and colleagues · read 28 Sept 2026
  5. repoLanguage Model Evaluation Harness, EleutherAI · read 28 Sept 2026