Data labeling
Data labeling attaches the answer a supervised model should learn to raw examples such as images, text, audio or table rows.
A label is the answer or value a supervised model is trained to predict from an example’s features. In a photo, for example, a person might name each object and trace its outline and position. Labels can also come from software or heuristic labeling functions rather than direct human judgment.
The work starts before annotation. A team defines the task, the allowed labels and the evidence annotators should use. Review then looks for unclear rules and disagreements. The approved result pairs each example with a target the model can compare against its prediction.
The resulting dataset should record its collection and labeling process so later users can judge whether it fits their task. A big pile of tags is not automatically trustworthy; a proxy label only approximates the real target, and raters can err or disagree, especially on value judgments.
Supervised models need examples paired with answers, but raw data arrives without those answers.
Follow one street photo from raw file to reviewed training example.
- 1 · defineWrite a label scheme and examples that state what counts as each class.
- 2 · assignGive the raw example and the same instructions to one or more annotators.
- 3 · reviewCompare the answers, look closely where labelers differ, and add clearer instructions when the guide caused the mix-up.
- 4 · storeSave the approved label beside the example and document how the dataset was created.
Clicking a tag is easy. Getting every labeler to use it the same way is the real work.
| Who | What they ask | What it works with |
|---|---|---|
| Vision team | “Which pixels belong to the pedestrian?” | Polygon and class attached to each image |
| Support team | “Is this message billing, technical or account help?” | Intent attached to each message |
| Speech team | “What words were spoken in this clip?” | Transcript aligned to the audio |
| Safety team | “Does this output break the written policy?” | Reviewer decision and rationale |
- Labeled examples give supervised training a target to predict.
- Image labels can say what each object is, trace its outline and mark where it sits.
- Multiple annotations and agreement checks can help assess whether an annotation task is reproducible.
- A stand-in (proxy) label only gets close to the outcome a team really cares about.
- Labelers agreeing with each other is a hint that a label is right, not proof of it.
- Weak supervision can create probabilistic labels without hand-labeling every example, but nobody knows up front how accurate its rules are or how much they overlap.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsDatasets: Labels, Google for Developers · read 27 Sept 2026
- docsSupervised Learning, Google for Developers · read 27 Sept 2026
- paperLabelMe: a database and web-based tool for image annotation, Russell et al. · read 27 Sept 2026
- paperSnorkel: Rapid Training Data Creation with Weak Supervision, Ratner et al. · read 27 Sept 2026
- paperDatasheets for Datasets, Gebru et al. · read 27 Sept 2026
- paperValidity, Agreement, Consensuality and Annotated Data Quality, Baledent et al., LREC 2022 · read 27 Sept 2026