Principal component analysis (PCA)
PCA finds the directions in which data spreads out most, then keeps only the first few, so many features become a few new ones.
Principal component analysis, or PCA, shrinks a table with many columns into one with a few new columns while losing as little information as possible. The new columns are called principal components. Each one is a weighted sum of the original features. The first carries as much variance as it can, each next one carries as much as is left, and no two of them are correlated. PCA is unsupervised: it uses only the features, never a label to predict. It dates back to work by Pearson in 1901 and Hotelling in 1933. More than a century on, it remains a standard way to make data smaller or to draw a picture of it.
The work happens in the covariance matrix, which records how each pair of features varies around their means together. Its top eigenvector is the first component. Seen another way, the first component is the best-fitting line through the points: it makes the average squared gap between each point and the line as small as possible. The same components can be found from the singular value decomposition (SVD) of the centred data, and scikit-learn’s PCA works that way. It subtracts the mean for you but leaves each feature’s units alone, so standardising is a job you do first.
How many components to keep is a judgement call. One rule of thumb keeps adding components until they cover about 70 percent of the variance, but picking that number is subjective. Another looks for an elbow in a scree plot, the point where each extra component adds little. No formula settles it for every data set. In scikit-learn’s PCA with the full solver, a number between 0 and 1 keeps just enough components to explain more than that share of the variance.
A textbook data set gives arrest rates for three crimes in each of the 50 US states. It adds a fourth measure, how much of each state’s population is urban. After standardising, the first component explained 62.0 percent of the variance and the second 24.7 percent, so a two-dimensional plot kept almost 87 percent. The first component weighted the three crime rates about equally, so it works as one score for how much serious crime a state has. PCA is linear, and manifold learning methods try to extend it to curved structure. t-SNE is one of those methods, and UMAP can draw similar maps of the data and can also shrink it along curves for other uses.
Data with many features is hard to look at, full of repeats and slow to work with.
Follow six points from two features down to one.
- 1 · centreSubtract each feature's mean, and usually divide by its standard deviation so no feature dominates just because of its units.
- 2 · findCompute the covariance matrix and take its eigenvector with the largest eigenvalue; that direction of greatest spread is the first principal component.
- 3 · rankEach later component is at right angles to the earlier ones, and its eigenvalue is the variance it carries.
- 4 · projectEach point's position along the kept components becomes its new, shorter list of numbers, called its scores.
PCA keeps the directions where the data varies most and drops the rest.
| Who | What they ask | What it works with |
|---|---|---|
| Data scientist | “Can I cut 50 correlated sensor readings down to a handful before training?” | The sensor columns, centred and scaled |
| Biologist | “Which samples look alike across thousands of gene measurements?” | A plot of the first two component scores |
| Computer vision student | “Can I describe a face with far fewer numbers than it has pixels?” | Face images, reduced to eigenfaces |
| Analyst | “How many components do I need to keep most of the variance?” | The explained variance ratio of each component |
- Replaces many correlated features with a few uncorrelated ones.
- Gives a two- or three-dimensional view of wide data for plotting.
- Shrinks inputs before training; a scikit-learn face example reduces 1,850 pixel features to 150 components.
- Dropping the later components often trims noise, because the real pattern in a data set tends to crowd into the first few.
- It only finds straight directions, so it can miss curved structure that non-linear methods such as t-SNE and UMAP aim to keep.
- Each component usually mixes all the original features, which can make it hard to interpret.
- Results depend on units; unscaled, the feature with the biggest variance takes over the first component.
- A few outliers can pull the components off course.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperPrincipal component analysis: a review and recent developments, Royal Society, Philosophical Transactions A (Jolliffe and Cadima, 2016) · read 27 Sept 2026
- paperAn Introduction to Statistical Learning (chapter 10, unsupervised learning), Springer (James, Witten, Hastie and Tibshirani) · read 27 Sept 2026
- paperMathematics for Machine Learning, chapter 10: Dimensionality Reduction with Principal Component Analysis, Cambridge University Press (Deisenroth, Faisal and Ong) · read 27 Sept 2026
- official6.5.5. Principal Components, NIST/SEMATECH e-Handbook of Statistical Methods · read 27 Sept 2026
- docsDecomposing signals in components (matrix factorization problems), scikit-learn · read 27 Sept 2026
- docsPCA (sklearn.decomposition.PCA), scikit-learn · read 27 Sept 2026
- docsFaces recognition example using eigenfaces and kernel approximation, scikit-learn · read 27 Sept 2026
- docsManifold learning, scikit-learn · read 27 Sept 2026
- docsUMAP: Uniform Manifold Approximation and Projection for Dimension Reduction, UMAP (umap-learn documentation) · read 27 Sept 2026