Unsupervised learning
Unsupervised learning trains a model on examples that have no answers attached, so it finds structure on its own, such as groups of similar items or points that do not fit.
Unsupervised learning is machine learning without an answer key. The model receives examples that have features, the details it can measure, but no label saying what the right answer is. Its task is to find structure on its own: which examples resemble each other, which details tend to go together, and which examples sit far from everything else.
The most common version is clustering, which sorts similar examples into groups. A popular method, k-means, drops a few centre points into the data, puts every example with its nearest centre, moves each centre to the middle of its group and repeats until the centres settle. The name comes from a paper by J. MacQueen, who described splitting data into k sets and suggested uses such as grouping similar items. You choose how many groups to look for. The model never learns what a group means, so a person looks at the results and names them.
Google Cloud’s BigQuery documentation gives a concrete case. It sorts London bike-hire stations into four groups using ride length, trips per day and distance from the city centre. One group turns out to be busy central stations; another is suburban stations with longer trips. No one told the model those kinds of station existed.
Clustering is not the only job. Dimensionality reduction squeezes many features into a few that keep most of the variation; principal component analysis (PCA) is the classic method, and t-SNE draws high-dimensional data as a flat map. Anomaly detection learns where ordinary data sits and flags what lies far away. Compared with supervised learning, which predicts known kinds of answer, unsupervised learning describes data, and without an answer key its results are harder to check.
Lots of data comes with no answers attached.
Follow k-means as it groups points that carry no labels.
- 1 · gatherCollect examples that have features but no label, and choose which features count when comparing them.
- 2 · startPick how many groups to look for, k, and place k starting centres.
- 3 · assignPut every example in the group of its nearest centre.
- 4 · updateMove each centre to the average of its group, then repeat assign and update until the centres stop moving.
- 5 · interpretA person looks at what each group has in common and gives it a name.
Nothing in the data says what the groups are. The model finds similarity; people supply the meaning.
| Who | What they ask | What it works with |
|---|---|---|
| Bike-hire operator | “Which of our stations behave alike?” | Station ride lengths, daily trips and distance from the centre |
| Music service | “Which songs sound similar enough to recommend together?” | Properties of each track, with no genre supplied |
| Fraud team | “Which transactions look unlike all the others?” | Payments scored by how far they sit from normal patterns |
| Data analyst | “Can I see this 50-column table on one chart?” | Features squeezed into two dimensions for plotting |
- Finds groups of similar examples when nobody has labelled the data.
- Shrinks many features into a few, which makes data easier to plot and faster to process.
- Flags unusual examples, such as a sudden spike in a time series, without a list of past anomalies.
- Its output can feed other systems, such as a recommendation service.
- It does not say what a group means; a person has to interpret and name it.
- Methods like k-means need you to choose the number of groups in advance, and results depend on where the centres start.
- There is usually no answer key to score it against, so checking quality is harder than in supervised learning.
- Distance-based grouping struggles with very many features, uneven group sizes and outliers.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsWhat is Machine Learning? (Introduction to Machine Learning), Google for Developers · read 27 Sept 2026
- docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
- docsWhat is clustering? (Clustering course), Google for Developers · read 27 Sept 2026
- docsAdvantages and disadvantages of k-means (Clustering course), Google for Developers · read 27 Sept 2026
- docs2.3. Clustering (User Guide), scikit-learn · read 27 Sept 2026
- docs2.5. Decomposing signals in components (User Guide), scikit-learn · read 27 Sept 2026
- docs2.7. Novelty and Outlier Detection (User Guide), scikit-learn · read 27 Sept 2026
- docsCreate a k-means model to cluster London bicycle hires dataset, Google Cloud · read 27 Sept 2026
- docsK-Means Algorithm (Amazon SageMaker AI Developer Guide), Amazon Web Services · read 27 Sept 2026
- docsRandom Cut Forest (RCF) Algorithm (Amazon SageMaker AI Developer Guide), Amazon Web Services · read 27 Sept 2026
- paperVisualizing Data using t-SNE, Journal of Machine Learning Research (van der Maaten and Hinton, 2008) · read 27 Sept 2026
- paperSome methods for classification and analysis of multivariate observations, J. MacQueen, Fifth Berkeley Symposium on Mathematical Statistics and Probability (1967), via Project Euclid · read 27 Sept 2026