Concepts

Unsupervised learning

5 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Unsupervised learning trains a model on examples that have no answers attached, so it finds structure on its own, such as groups of similar items or points that do not fit.

1 · What it is

Unsupervised learning is machine learning without an answer key. The model receives examples that have features, the details it can measure, but no label saying what the right answer is. Its task is to find structure on its own: which examples resemble each other, which details tend to go together, and which examples sit far from everything else.

The most common version is clustering, which sorts similar examples into groups. A popular method, k-means, drops a few centre points into the data, puts every example with its nearest centre, moves each centre to the middle of its group and repeats until the centres settle. The name comes from a paper by J. MacQueen, who described splitting data into k sets and suggested uses such as grouping similar items. You choose how many groups to look for. The model never learns what a group means, so a person looks at the results and names them.

Google Cloud’s BigQuery documentation gives a concrete case. It sorts London bike-hire stations into four groups using ride length, trips per day and distance from the city centre. One group turns out to be busy central stations; another is suburban stations with longer trips. No one told the model those kinds of station existed.

Clustering is not the only job. Dimensionality reduction squeezes many features into a few that keep most of the variation; principal component analysis (PCA) is the classic method, and t-SNE draws high-dimensional data as a flat map. Anomaly detection learns where ordinary data sits and flags what lies far away. Compared with supervised learning, which predicts known kinds of answer, unsupervised learning describes data, and without an answer key its results are harder to check.

2 · Why it exists

Lots of data comes with no answers attached.

Labels are missingSometimes useful labels are scarce or missing entirely, so there is no right answer for a model to copy.
The categories are unknownSometimes you do not yet know what groups exist, so there is nothing to put in a label column.
Too many detailsData with dozens of features cannot be drawn on a flat chart as it is, and distances between examples grow less informative.
3 · How it works

Follow k-means as it groups points that carry no labels.

1 · INPUT: STATIONS, NO LABELS K-MEANS, K = 3 OUTPUT: 3 UNNAMED GROUPS trips per day distance from centre 2 · start pick 3 random points as centres 3 · assign each point joins its nearest centre 4 · update move each centre to its group mean repeat repeat 3 and 4 until the centres stop moving done trips per day distance from centre group 1 group 2 group 3 No labels go in. k-means groups points by distance alone, so every point ends up with its nearest centre. A person names the groups afterwards, for example busy central stations or suburban ones with longer trips.
k-means only measures distance. Deciding that group 1 means busy central stations is a human step.
  1. 1 · gatherCollect examples that have features but no label, and choose which features count when comparing them.
  2. 2 · startPick how many groups to look for, k, and place k starting centres.
  3. 3 · assignPut every example in the group of its nearest centre.
  4. 4 · updateMove each centre to the average of its group, then repeat assign and update until the centres stop moving.
  5. 5 · interpretA person looks at what each group has in common and gives it a name.

Nothing in the data says what the groups are. The model finds similarity; people supply the meaning.

4 · Where it's used
WhoWhat they askWhat it works with
Bike-hire operator“Which of our stations behave alike?”Station ride lengths, daily trips and distance from the centre
Music service“Which songs sound similar enough to recommend together?”Properties of each track, with no genre supplied
Fraud team“Which transactions look unlike all the others?”Payments scored by how far they sit from normal patterns
Data analyst“Can I see this 50-column table on one chart?”Features squeezed into two dimensions for plotting
5 · What it solves, and what it doesn't
solves
  • Finds groups of similar examples when nobody has labelled the data.
  • Shrinks many features into a few, which makes data easier to plot and faster to process.
  • Flags unusual examples, such as a sudden spike in a time series, without a list of past anomalies.
  • Its output can feed other systems, such as a recommendation service.
doesn't solve
  • It does not say what a group means; a person has to interpret and name it.
  • Methods like k-means need you to choose the number of groups in advance, and results depend on where the centres start.
  • There is usually no answer key to score it against, so checking quality is harder than in supervised learning.
  • Distance-based grouping struggles with very many features, uneven group sizes and outliers.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsWhat is Machine Learning? (Introduction to Machine Learning), Google for Developers · read 27 Sept 2026
  2. docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
  3. docsWhat is clustering? (Clustering course), Google for Developers · read 27 Sept 2026
  4. docsAdvantages and disadvantages of k-means (Clustering course), Google for Developers · read 27 Sept 2026
  5. docs2.3. Clustering (User Guide), scikit-learn · read 27 Sept 2026
  6. docs2.5. Decomposing signals in components (User Guide), scikit-learn · read 27 Sept 2026
  7. docs2.7. Novelty and Outlier Detection (User Guide), scikit-learn · read 27 Sept 2026
  8. docsCreate a k-means model to cluster London bicycle hires dataset, Google Cloud · read 27 Sept 2026
  9. docsK-Means Algorithm (Amazon SageMaker AI Developer Guide), Amazon Web Services · read 27 Sept 2026
  10. docsRandom Cut Forest (RCF) Algorithm (Amazon SageMaker AI Developer Guide), Amazon Web Services · read 27 Sept 2026
  11. paperVisualizing Data using t-SNE, Journal of Machine Learning Research (van der Maaten and Hinton, 2008) · read 27 Sept 2026
  12. paperSome methods for classification and analysis of multivariate observations, J. MacQueen, Fifth Berkeley Symposium on Mathematical Statistics and Probability (1967), via Project Euclid · read 27 Sept 2026