Ground truth
Ground truth is the best available account of what really happened, the reality that a model's predictions are checked against.
Ground truth is what actually happened: the real answer that a model’s predictions are measured against. If a model predicts whether a first-year university student will graduate within six years, the ground truth is whether that student really did. A label is the answer attached to one example. Ground truth is the reality that label is meant to record, and a label is only an estimate of it.
How ground truth is found depends on the question. Sometimes you wait for the outcome, as with graduation. Sometimes an instrument measures it. In medicine, a new test is judged against a reference standard, the best available way of deciding whether a patient has a condition, which may combine several tests with follow-up over time. When the answer is a matter of opinion, several people rate each example and their answers are combined; Amazon’s labelling service estimates the true class from each worker’s answers. Satellite scientists do something similar and call it ground validation, or ground truthing: they check what a satellite reports against rain gauges and other measurements taken on the ground.
Scoring is then counting. Each prediction is compared with its ground truth and lands in one of four boxes, which together make a table called a confusion matrix. Accuracy is the right answers divided by all answers.
The catch is that ground truth is the best available account, not a perfect one. Records contain mistakes, instruments differ, and people disagree. A 2021 study estimated that at least 3.3% of test labels across 10 widely used datasets were wrong, including at least 6% of the ImageNet validation set. That can reorder models: on ImageNet with corrected labels, the smaller ResNet-18 overtook ResNet-50 once the share of mislabelled test examples rose by just 6%.
A model's score means nothing without something real to check it against.
Follow one test through scoring.
- 1 · predictThe model makes a prediction for every test case, such as tumour or no tumour.
- 2 · establishGround truth for the same cases comes from a separate source, such as an expert's diagnosis, a later outcome or an instrument reading.
- 3 · compareEach prediction is matched against its ground truth and counted as right or wrong.
- 4 · tallyThe counts fill a table called a confusion matrix, which gives scores such as accuracy.
A score measures agreement with ground truth, so mistakes in the ground truth become mistakes in the score.
| Who | What they ask | What it works with |
|---|---|---|
| Hospital AI team | “Does our scan reader agree with the biopsy results?” | The model's readings against the reference standard for each patient |
| Satellite rainfall team | “Does the satellite's rain estimate match what fell on the ground?” | Satellite estimates against rain gauges and ground radar |
| Benchmark maintainers | “How many of our test answers are actually wrong?” | Test-set labels flagged for a second human check |
| Email spam team | “How many real spam emails slipped into the inbox?” | Filter decisions against emails known to be spam |
- Gives every prediction something real to be scored against.
- Turns evaluation into counting right and wrong answers in a confusion matrix.
- Shows which kinds of mistake a model makes, not only how many.
- Lets people check a remote reading, such as a satellite's, against a direct one on the ground.
- It is not guaranteed to be true; records, instruments and raters make errors.
- Some ground truth takes years to arrive, as with a six-year graduation outcome.
- Where no reference standard exists, agreement with another test is not true accuracy.
- Asking more raters can make labels more accurate, but it also costs more.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
- docsClassification: Thresholds and the confusion matrix (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docs3.4. Metrics and scoring: quantifying the quality of predictions, scikit-learn · read 27 Sept 2026
- paperPervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks, NeurIPS 2021 Datasets and Benchmarks Track (via arXiv) · read 27 Sept 2026
- officialStatistical Guidance on Reporting Results from Studies Evaluating Diagnostic Tests, U.S. Food and Drug Administration · read 27 Sept 2026
- docsAnnotation consolidation (Amazon SageMaker Ground Truth), Amazon Web Services · read 27 Sept 2026
- officialGround Validation and OLYMPEX Webquest, NASA Global Precipitation Measurement mission · read 27 Sept 2026
- paperField calibration and validation of remote-sensing surveys, U.S. Geological Survey · read 27 Sept 2026