Labels
A label is the answer attached to a training example, such as a category, a number or a ranking, that a supervised model learns to predict.
A label is the answer attached to a training example: the part a supervised model is trying to learn to predict. The rest of the example, its features, is what the model gets to look at. Picture a table of house sales. Bedrooms, bathrooms and the house’s age are the features, and the sale price is the label. A model trains on labelled examples, then predicts labels for new examples that have none. Labels are meant to record the ground truth, what really happened, but people and programs write them down, so they can be wrong.
Labels come in several shapes. A label can be a category: ImageNet sorts its photos into concepts, and people check and label them. It can also be a number, like that sale price, or a drawing: the 2014 paper that introduced the COCO image dataset reported 2.5 million objects, each traced with its own outline, across 328,000 photos. For chat models a label can be a ranking. To train InstructGPT, raters put four to nine answers to the same prompt in order from best to worst, and a reward model learned to predict their choice.
Sometimes the answer you want was never recorded. A business that wants to reach bicycle owners may have no column saying who owns one, only who recently bought one. That stand-in is a proxy label. It is a fairly good one, but gifts break the link, and the model is only as useful as that link. Labels come from raters, who are usually paid, from users’ own actions such as tagging photos, or from another model. Labels made by people are often called gold labels; labels made by a model are called silver.
Gold does not mean flawless. Researchers who audited ten popular benchmark test sets, covering images, text and audio, estimated that at least 3.3% of the labels were wrong on average, and for ImageNet’s validation set the figure was at least 6%. Even screened raters disagree: InstructGPT’s raters, picked partly for how well they matched the researchers, agreed with each other about 73% of the time. So teams can send each example to several raters, measure how often they agree, merge the answers into one label and spot-check raters against a sample they labelled themselves.
A model can only be as right as the answers it studies.
Follow one example from raw item to trusted label.
- 1 · defineWrite step-by-step instructions and show raters the full set of labels they can choose from.
- 2 · labelSeveral raters label the same example, each one working on it separately.
- 3 · agreeCompare their answers; if enough match, that answer is kept, and if not the example goes to review.
- 4 · checkTest the raters by labelling a sample yourself and measuring how often their answers match yours.
- 5 · trainThe example and its single agreed label join the training data.
Agreement is evidence, not proof. Even when every rater agrees, the label can still be wrong.
| Who | What they ask | What it works with |
|---|---|---|
| Self-driving car team | “Where exactly is each pedestrian in this camera frame?” | Boxes and outlines drawn around every object |
| Chatbot lab | “Which of these four replies is the most helpful?” | Rankings of model answers made by trained raters |
| Hospital research group | “Does this scan show a tumour?” | Scans read by trained specialists |
| Photo app | “Which of these photos show the user's dog?” | Tags that users add to their own photos |
- Gives a supervised model a target to learn from and a score to be measured against.
- Lets raters put several answers in order, such as chatbot replies from best to worst, instead of naming one right answer.
- Using more raters per example can make labels more accurate.
- Low agreement between raters can be a sign that the instructions need work.
- Labels are written by people or programs, so even famous test sets contain wrong ones.
- A proxy label is always a compromise; the model learns the stand-in, not the real target.
- More raters per example means higher labelling costs.
- Each example usually keeps just one label even when its raters disagreed, so those disagreements are worth going back to review.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
- docsDatasets: Labels (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsCategorical data: Common issues (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docsData Collection + Evaluation (People + AI Guidebook), Google People + AI Research (PAIR) · read 27 Sept 2026
- paperPervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks, arXiv (MIT, Northcutt et al.) · read 27 Sept 2026
- paperTraining language models to follow instructions with human feedback, arXiv (OpenAI, Ouyang et al.) · read 27 Sept 2026
- paperMicrosoft COCO: Common Objects in Context, arXiv (Microsoft and academic partners, Lin et al.) · read 27 Sept 2026
- docsAnnotation consolidation (Amazon SageMaker Ground Truth), Amazon Web Services · read 27 Sept 2026
- officialAbout ImageNet, ImageNet (Stanford University and Princeton University) · read 27 Sept 2026