Clustering
Clustering sorts data that has no labels into groups of items that resemble each other.
Clustering is a form of unsupervised learning; nobody tells it which group any example belongs to. If the examples already carried group labels, the same sorting job would be called classification instead. When the algorithm finishes, every group comes back tagged with its own cluster ID. Nothing tells you whether those groups are right, so a person has to judge them, looking both at whole groups and at single examples.
Among methods built around centroids, k-means is the most popular. A centroid is simply the average position of the points in its group. You choose k, the number of groups, before it runs. The starting centroids are dropped at random. Each example then joins whichever centroid is closest. Next, each centroid moves to the average spot of the examples that joined it. Those two moves repeat until no example switches group. Because the start is random, two runs on the same data can end with quite different groups.
How you measure similarity is a decision that depends on knowing your data well. Raw features first need normalizing, scaling and transforming, just as in any other machine learning project. With two features you can eyeball which points belong together; with many, weighing and comparing them stops being obvious.
Other algorithms bake in other assumptions, and each one fits a certain kind of data best. Hierarchical clustering produces a tree, and cutting that tree at a chosen level gives the number of clusters you want. Density-based methods join up crowded regions, so a group can take any shape, and stray outliers stay unassigned. DBSCAN, described in a 1996 paper, is a clustering algorithm built on this density-based view of what a cluster is. Density methods have weak spots too; they struggle when clusters differ in density or when the data has many dimensions.
Data with no labels can still hold groups worth finding.
Follow six customers through one k-means update.
- 1 · prepareClean and rescale the features, then confirm that distances between examples now reflect real similarity.
- 2 · seedChoose k, the number of clusters. Then place k starting centroids at random.
- 3 · assignSend every example to its nearest centroid.
- 4 · updateMove each centroid to the average position of the examples assigned to it.
- 5 · repeatReassign and update until no example switches group, or stop early on another rule when the dataset is large.
A cluster ID marks a discovered group; you still have to check whether that group makes sense.
| Who | What they ask | What it works with |
|---|---|---|
| Retail analyst | “Which customers behave similarly?” | Purchase frequency, value and product mix |
| Search team | “Which results cover the same topic?” | Document representations |
| Biologist | “Which species share hidden genetic similarities?” | Gene sequence data |
| Operations team | “Which events fall outside every normal group?” | Event feature vectors |
- Finds groups of similar items in data with no labels.
- One cluster ID can stand in for many features, which shrinks the data.
- K-means scales as O(nk), so it stays practical on large datasets.
- Density-based methods can find clusters of any shape and leave outliers out.
- Gives no ground truth to check the groups against.
- K-means requires the number of clusters to be chosen before fitting.
- Random starting centroids can make k-means give quite different groups on different runs.
- K-means assumes roughly round groups, so outliers and density-shaped clusters can mislead it.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsWhat is clustering?, Google for Developers · read 27 Sept 2026
- docsClustering workflow, Google for Developers · read 27 Sept 2026
- docsWhat is k-means clustering?, Google for Developers · read 27 Sept 2026
- docsClustering algorithms, Google for Developers · read 27 Sept 2026
- paperA Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise, Ester, Kriegel, Sander and Xu, KDD-96 (AAAI) · read 27 Sept 2026