Concepts

Training data

5 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Training data is the set of examples a machine learning model studies to learn its patterns, kept apart from the examples used to test it.

1 · What it is

Training data is the collection of examples a machine learning model learns from. It reads the examples, often many times over, and slowly adjusts its internal numbers, called parameters, until its outputs fit them. Whatever is in the data, good or bad, shapes the model, and a weak dataset caps how good the model can get.

Builders rarely train on everything they collect. They split it into three separate parts: a training set the model studies, a validation set for checks along the way, and a test set held back for a final check. Each example belongs to one part only. If copies of test examples slip into training, the model is being graded on questions it has already seen, and its score looks better than it should. For language models this is called contamination, because test questions posted online can end up in crawled training text.

Large language models need a lot of data. GPT-3’s training mix drew 60 per cent from a filtered copy of Common Crawl, a free archive that now holds over 300 billion web pages, with the rest from curated web text, books and Wikipedia. The team filtered out low-quality pages and removed near-duplicates first. Models tuned to follow instructions, such as OpenAI’s InstructGPT, add a second kind of data, example answers written by people. Research on the Chinchilla model found that many large models had been trained on too little data for their size.

Because the data shapes the model this much, researchers and lawmakers have asked for it to be written down and checked. “Datasheets for datasets” proposes a written record of why a dataset was made, what is in it and how it was collected. In the EU, the AI Act says the data behind high-risk AI systems must be of high quality, so the systems are less likely to treat groups of people unfairly. The law’s article on data adds that the training, validation and test sets must suit the intended purpose, represent it well enough, and be as complete and error-free as possible. After the 2026 AI Omnibus, the high-risk rules apply from 2 December 2027 for sensitive uses such as hiring and education. For AI built into products such as toys, they apply from 2 August 2028. Obligations for general-purpose model providers have applied since 2 August 2025. Makers of general-purpose models must publish an overview of the data they trained on, following a template from the European Commission.

2 · Why it exists

A model learns only what its examples show it.

Examples do the teachingA model reads examples and nudges its internal settings to fit them, so the examples decide what it learns.
Flaws get learned tooWrong labels, duplicates and slanted samples in the data end up in the model's behaviour.
Scores need unseen dataTesting a model on examples it already studied says little about how it will handle new ones.
3 · How it works

Follow data from its source to the model.

Every example lands in one set only. The model learns from the training set; the test set stays unseen until the end.
  1. 1 · collectGather examples from sources such as web crawls, books and reference text, or answers written by people.
  2. 2 · cleanFilter out low-quality material, fix or drop broken examples and remove duplicates.
  3. 3 · documentWrite down why the dataset exists, what is in it and how it was collected.
  4. 4 · splitDivide the examples into a training set, a validation set and a test set, with each example in only one.
  5. 5 · trainThe model learns from the training set, the validation set guides tweaks, and the test set gives the final check.

Test examples are never trained on. If copies slip into training, the final score is too good to trust.

4 · Where it's used
WhoWhat they askWhat it works with
Language model lab“Which crawled pages are worth keeping, and which are near-copies of each other?”Raw web text before filtering and deduplication
Hospital imaging team“Do our scans cover every patient group the tool will be used on?”Labelled medical images and who they come from
Online shop“Has buying behaviour changed since we last trained the model?”Recent purchase records compared with the training period
Compliance officer“Can we show where our training data came from and how it was checked?”Dataset documentation and data-governance records
5 · What it solves, and what it doesn't
solves
  • Lets a model pick up patterns by tuning its parameters to fit many examples.
  • A held-back test set gives a fair estimate of how the model does on new data.
  • More examples, and more varied ones, make overfitting less likely. Overfitting means memorising the examples instead of generalising.
  • Written documentation helps the people who reuse a dataset understand what it is.
doesn't solve
  • More data does not cure bad data; label errors, bad values and duplicates still hurt.
  • Data collected once goes stale when the world it describes changes.
  • Slanted examples can lead a model to repeat stereotypes and prejudice.
  • Models can memorise rare training text, including personal details, and repeat it word for word.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsDatasets: Dividing the original dataset (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  2. docsDatasets: Data characteristics (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  3. docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
  4. docsProduction ML systems: Static versus dynamic training (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
  5. paperLanguage Models are Few-Shot Learners, arXiv (OpenAI, Brown et al.) · read 27 Sept 2026
  6. paperTraining Compute-Optimal Large Language Models, arXiv (DeepMind, Hoffmann et al.) · read 27 Sept 2026
  7. paperTraining language models to follow instructions with human feedback, arXiv (OpenAI, Ouyang et al.) · read 27 Sept 2026
  8. officialCommon Crawl - Open Repository of Web Crawl Data, Common Crawl Foundation · read 27 Sept 2026
  9. paperDatasheets for Datasets, arXiv (Gebru et al.; Communications of the ACM, 2021) · read 27 Sept 2026
  10. paperExtracting Training Data from Large Language Models, arXiv (Carlini et al.) · read 27 Sept 2026
  11. officialRegulation (EU) 2024/1689 (Artificial Intelligence Act), Article 10, European Union (EUR-Lex) · read 27 Sept 2026
  12. officialAI Act, European Commission · read 27 Sept 2026
  13. officialAI Omnibus enters into force, European Commission · read 27 Sept 2026