Concepts

Ground truth

5 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Ground truth is the best available account of what really happened, the reality that a model's predictions are checked against.

1 · What it is

Ground truth is what actually happened: the real answer that a model’s predictions are measured against. If a model predicts whether a first-year university student will graduate within six years, the ground truth is whether that student really did. A label is the answer attached to one example. Ground truth is the reality that label is meant to record, and a label is only an estimate of it.

How ground truth is found depends on the question. Sometimes you wait for the outcome, as with graduation. Sometimes an instrument measures it. In medicine, a new test is judged against a reference standard, the best available way of deciding whether a patient has a condition, which may combine several tests with follow-up over time. When the answer is a matter of opinion, several people rate each example and their answers are combined; Amazon’s labelling service estimates the true class from each worker’s answers. Satellite scientists do something similar and call it ground validation, or ground truthing: they check what a satellite reports against rain gauges and other measurements taken on the ground.

Scoring is then counting. Each prediction is compared with its ground truth and lands in one of four boxes, which together make a table called a confusion matrix. Accuracy is the right answers divided by all answers.

The catch is that ground truth is the best available account, not a perfect one. Records contain mistakes, instruments differ, and people disagree. A 2021 study estimated that at least 3.3% of test labels across 10 widely used datasets were wrong, including at least 6% of the ImageNet validation set. That can reorder models: on ImageNet with corrected labels, the smaller ResNet-18 overtook ResNet-50 once the share of mislabelled test examples rose by just 6%.

2 · Why it exists

A model's score means nothing without something real to check it against.

Predictions are not realityA model's output is a guess, often a probability, and a guess is not what actually happened.
Truth can arrive lateFor some questions the real answer only shows up years later, such as whether a student graduates within six years.
Truth can be wrongRecords, instruments and human raters all make mistakes, so ground truth is not always fully true.
3 · How it works

Follow one test through scoring.

Counts from the tumour example in Google's Machine Learning Glossary; the 3.3% figure is from Northcutt and colleagues' 2021 study of benchmark test sets.
  1. 1 · predictThe model makes a prediction for every test case, such as tumour or no tumour.
  2. 2 · establishGround truth for the same cases comes from a separate source, such as an expert's diagnosis, a later outcome or an instrument reading.
  3. 3 · compareEach prediction is matched against its ground truth and counted as right or wrong.
  4. 4 · tallyThe counts fill a table called a confusion matrix, which gives scores such as accuracy.

A score measures agreement with ground truth, so mistakes in the ground truth become mistakes in the score.

4 · Where it's used
WhoWhat they askWhat it works with
Hospital AI team“Does our scan reader agree with the biopsy results?”The model's readings against the reference standard for each patient
Satellite rainfall team“Does the satellite's rain estimate match what fell on the ground?”Satellite estimates against rain gauges and ground radar
Benchmark maintainers“How many of our test answers are actually wrong?”Test-set labels flagged for a second human check
Email spam team“How many real spam emails slipped into the inbox?”Filter decisions against emails known to be spam
5 · What it solves, and what it doesn't
solves
  • Gives every prediction something real to be scored against.
  • Turns evaluation into counting right and wrong answers in a confusion matrix.
  • Shows which kinds of mistake a model makes, not only how many.
  • Lets people check a remote reading, such as a satellite's, against a direct one on the ground.
doesn't solve
  • It is not guaranteed to be true; records, instruments and raters make errors.
  • Some ground truth takes years to arrive, as with a six-year graduation outcome.
  • Where no reference standard exists, agreement with another test is not true accuracy.
  • Asking more raters can make labels more accurate, but it also costs more.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
  2. docsClassification: Thresholds and the confusion matrix (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  3. docs3.4. Metrics and scoring: quantifying the quality of predictions, scikit-learn · read 27 Sept 2026
  4. paperPervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks, NeurIPS 2021 Datasets and Benchmarks Track (via arXiv) · read 27 Sept 2026
  5. officialStatistical Guidance on Reporting Results from Studies Evaluating Diagnostic Tests, U.S. Food and Drug Administration · read 27 Sept 2026
  6. docsAnnotation consolidation (Amazon SageMaker Ground Truth), Amazon Web Services · read 27 Sept 2026
  7. officialGround Validation and OLYMPEX Webquest, NASA Global Precipitation Measurement mission · read 27 Sept 2026
  8. paperField calibration and validation of remote-sensing surveys, U.S. Geological Survey · read 27 Sept 2026