Supervised learning
Supervised learning trains a model on examples that already carry the right answer, called a label, so it can predict that answer for new examples that do not.
Supervised learning trains a model on examples whose answers are already known. Each example has features, the details the model can see, and a label, the answer you want it to learn to give. An email’s links and wording are features; “spam” or “not spam” is the label.
Training is a loop of guessing and correcting. The model predicts a label from the features, the prediction is compared with the real label, and the model adjusts itself to shrink the gap. Over many examples, it settles on a relationship between features and labels that works well on average.
The labels are what make it supervised: they act as the teacher’s answer key. That is also its main cost. Someone has to supply those answers, often by hand, and weak examples or sloppy labels lead to a weak model. ImageNet, a research dataset of photos labelled by people, shows the scale involved: a well-known 2012 model trained on 1.3 million of its labelled images to recognise 1,000 kinds of object. When the answer is a number, such as rainfall, the task is called regression; when it is a category, such as spam, it is called classification.
A model can learn to answer a question by studying many past examples that already have the answer.
Learn from answered examples, then answer new ones.
- 1 · collectGather examples, each with features (what the model sees) and a label (the answer you want it to predict).
- 2 · compareThe model predicts a label from the features, and the gap between its guess and the real label is the loss.
- 3 · adjustThe model updates itself to shrink that loss, repeating across the whole dataset, often several times.
- 4 · checkTest it on labelled examples it has not trained on, showing it only the features.
- 5 · predictUse the trained model on new, unlabelled examples.
The labels are the supervision. Without them, the model has nothing to be scored against.
| Who | What they ask | What it works with |
|---|---|---|
| Email provider | “Is this message spam?” | Past messages marked spam or not spam |
| Weather app | “How much rain will fall?” | Past readings with the rainfall that followed |
| Property site | “What is this house worth?” | Past sales with their final prices |
| Photo service | “Does this picture contain a cat?” | Images tagged with what they show |
- Learns the link between features and answers straight from labelled examples.
- Handles both numbers (regression, such as rainfall) and categories (classification, such as spam).
- Its accuracy can be measured directly, because the right answers are known for test examples.
- It needs labelled data, and paying people to label it costs money and invites mistakes.
- Its predictions can be no better than its examples, so a small or narrow dataset leads to weak answers.
- When the label is a substitute for the real target, a weak tie between the two limits what the model can deliver.
- Extra inputs are not free wins, because an input with no real effect on the answer adds nothing.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsSupervised Learning (Introduction to Machine Learning), Google for Developers · read 27 Sept 2026
- docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
- docsWhat is Machine Learning? (Introduction to Machine Learning), Google for Developers · read 27 Sept 2026
- docsDatasets: Labels (Machine Learning Crash Course), Google for Developers · read 27 Sept 2026
- docs1. Supervised learning (User Guide), scikit-learn · read 27 Sept 2026
- officialAbout ImageNet, ImageNet · read 27 Sept 2026
- paperImageNet Classification with Deep Convolutional Neural Networks, NeurIPS 2012 (Krizhevsky, Sutskever and Hinton) · read 27 Sept 2026