Linear algebra for machine learning
The small part of linear algebra that models use every day, vectors, matrices, their products and their sizes, to hold data, run layers and learn.
Linear algebra is the study of vectors and matrices. Machine learning uses a small, well-defined slice of it: a deep learning textbook’s linear algebra chapter leaves out whole topics that are not essential for deep learning. MIT’s 18.065 course teaches it alongside probability, statistics and optimisation as a route into deep learning.
Machine learning represents data as vectors. A feature vector holds the feature values of one example, and an embedding vector is a list of numbers learned during training, often in an embedding layer. A table of examples is a matrix with one row per example and one column per feature. A layer’s weights form a matrix as well; PyTorch’s Linear layer stores them with shape out_features by in_features.
The workhorse is multiplication. A single neuron multiplies each input by its weight and adds the products. That sum of products is a dot product. Doing it for every example in a batch and every output at once is one matrix product. PyTorch’s Linear layer writes this as y = xAᵀ + b, and accepts inputs with any number of leading batch dimensions. The small T marks the transpose, which mirrors a matrix across its main diagonal so rows become columns. Two matrices can be multiplied only when their inner sizes match. Transformer attention computes the dot products of a query with every key. Many queries are packed into a matrix and handled together. The authors note this is fast because it runs on highly optimised matrix multiplication code. In NumPy, such routines rely on the BLAS and LAPACK libraries.
A norm measures the size of a vector. The usual one, the L2 norm, is the straight-line distance from the origin to the vector’s tip. L2 regularisation penalises the sum of the squared weights. This pulls the weights towards zero.
A matrix decomposition breaks a matrix into simpler factors, much as a number is factored. The singular value decomposition is a central example. Principal component analysis uses the SVD of centred data to project it onto fewer dimensions. The directions PCA keeps are eigenvectors of the data’s covariance matrix with the largest eigenvalues. Gradient descent moves the weights a small step against the gradient.
Machine learning needs one compact way to hold many numbers and apply the same sum to all of them.
Follow one prediction through the operations it uses.
- 1 · representEach input becomes a vector, a fixed-order list of numbers such as the feature values of one example.
- 2 · stackA batch of inputs is stacked into a matrix, one row per example.
- 3 · multiplyA layer multiplies that matrix by a weight matrix and adds a bias, giving every output for every input in one product.
- 4 · compareA dot product turns two vectors into one score, as when attention compares a query with each key.
- 5 · updateTraining subtracts a scaled gradient from the weights, which is plain vector arithmetic.
A linear layer is one matrix product plus a bias, done for a whole batch at once.
| Who | What they ask | What it works with |
|---|---|---|
| Machine learning engineer | “Why does this layer reject my input tensor?” | The inner sizes of the input and weight matrices |
| Search team | “Which stored item scores highest against this query?” | Dot products between embedding vectors |
| Data scientist | “Can I shrink 100 features to 10 before training?” | PCA, computed with the SVD of the centred data |
| Model trainer | “Are the weights growing too large?” | The squared L2 norm of the weights, added as a penalty |
- Gives data, weights and outputs one shared form, vectors and matrices with known shapes.
- Writes a whole layer over a whole batch as one matrix product.
- Hands that product to tuned libraries such as BLAS and LAPACK.
- Supplies ready tools for size, similarity and compression, such as norms, dot products and the SVD.
- It is not all the maths; training also relies on vector calculus, and noisy data needs probability.
- Matrix products alone are linear, so networks add activation functions to learn curved relationships.
- A dot product is a simple score, and the transformer authors suggest a richer comparison may help.
- It does not pick good features; that remains separate work.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperMathematics for Machine Learning (book PDF), Cambridge University Press (Deisenroth, Faisal and Ong) · read 27 Sept 2026
- paperDeep Learning, chapter 2: Linear Algebra, MIT Press (Goodfellow, Bengio and Courville) · read 27 Sept 2026
- docsMatrix Methods in Data Analysis, Signal Processing, and Machine Learning (18.065), MIT OpenCourseWare · read 27 Sept 2026
- docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
- docsLinear (PyTorch 2.14 documentation), PyTorch · read 27 Sept 2026
- paperAttention Is All You Need, Vaswani et al., NeurIPS 2017 · read 27 Sept 2026
- docsLinear algebra (numpy.linalg), NumPy · read 27 Sept 2026
- docstorch.linalg (PyTorch 2.14 documentation), PyTorch · read 27 Sept 2026
- docssklearn.decomposition.PCA, scikit-learn · read 27 Sept 2026