Linear algebra for MLConcepts

Linear algebra for machine learning

4 min readbeginnerUpdated 28 Sept 2026
1 · In one line

The small part of linear algebra that models use every day, vectors, matrices, their products and their sizes, to hold data, run layers and learn.

1 · What it is

Linear algebra is the study of vectors and matrices. Machine learning uses a small, well-defined slice of it: a deep learning textbook’s linear algebra chapter leaves out whole topics that are not essential for deep learning. MIT’s 18.065 course teaches it alongside probability, statistics and optimisation as a route into deep learning.

Machine learning represents data as vectors. A feature vector holds the feature values of one example, and an embedding vector is a list of numbers learned during training, often in an embedding layer. A table of examples is a matrix with one row per example and one column per feature. A layer’s weights form a matrix as well; PyTorch’s Linear layer stores them with shape out_features by in_features.

The workhorse is multiplication. A single neuron multiplies each input by its weight and adds the products. That sum of products is a dot product. Doing it for every example in a batch and every output at once is one matrix product. PyTorch’s Linear layer writes this as y = xAᵀ + b, and accepts inputs with any number of leading batch dimensions. The small T marks the transpose, which mirrors a matrix across its main diagonal so rows become columns. Two matrices can be multiplied only when their inner sizes match. Transformer attention computes the dot products of a query with every key. Many queries are packed into a matrix and handled together. The authors note this is fast because it runs on highly optimised matrix multiplication code. In NumPy, such routines rely on the BLAS and LAPACK libraries.

A norm measures the size of a vector. The usual one, the L2 norm, is the straight-line distance from the origin to the vector’s tip. L2 regularisation penalises the sum of the squared weights. This pulls the weights towards zero.

A matrix decomposition breaks a matrix into simpler factors, much as a number is factored. The singular value decomposition is a central example. Principal component analysis uses the SVD of centred data to project it onto fewer dimensions. The directions PCA keeps are eigenvectors of the data’s covariance matrix with the largest eigenvalues. Gradient descent moves the weights a small step against the gradient.

2 · Why it exists

Machine learning needs one compact way to hold many numbers and apply the same sum to all of them.

Data needs a layoutA model can only compute with numbers, so every example has to become a list of numbers, and many examples a table.
Same sum, many timesA layer applies the same weights to every input, so code needs to push many inputs through one call.
Papers assume itMachine learning algorithms, deep learning above all, are described in the language of vectors and matrices.
3 · How it works

Follow one prediction through the operations it uses.

Illustrative shapes. The highlighted multiply does a whole layer for a whole batch in one step; the dashed loop is training.
  1. 1 · representEach input becomes a vector, a fixed-order list of numbers such as the feature values of one example.
  2. 2 · stackA batch of inputs is stacked into a matrix, one row per example.
  3. 3 · multiplyA layer multiplies that matrix by a weight matrix and adds a bias, giving every output for every input in one product.
  4. 4 · compareA dot product turns two vectors into one score, as when attention compares a query with each key.
  5. 5 · updateTraining subtracts a scaled gradient from the weights, which is plain vector arithmetic.

A linear layer is one matrix product plus a bias, done for a whole batch at once.

4 · Where it's used
WhoWhat they askWhat it works with
Machine learning engineer“Why does this layer reject my input tensor?”The inner sizes of the input and weight matrices
Search team“Which stored item scores highest against this query?”Dot products between embedding vectors
Data scientist“Can I shrink 100 features to 10 before training?”PCA, computed with the SVD of the centred data
Model trainer“Are the weights growing too large?”The squared L2 norm of the weights, added as a penalty
5 · What it solves, and what it doesn't
solves
  • Gives data, weights and outputs one shared form, vectors and matrices with known shapes.
  • Writes a whole layer over a whole batch as one matrix product.
  • Hands that product to tuned libraries such as BLAS and LAPACK.
  • Supplies ready tools for size, similarity and compression, such as norms, dot products and the SVD.
doesn't solve
  • It is not all the maths; training also relies on vector calculus, and noisy data needs probability.
  • Matrix products alone are linear, so networks add activation functions to learn curved relationships.
  • A dot product is a simple score, and the transformer authors suggest a richer comparison may help.
  • It does not pick good features; that remains separate work.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperMathematics for Machine Learning (book PDF), Cambridge University Press (Deisenroth, Faisal and Ong) · read 27 Sept 2026
  2. paperDeep Learning, chapter 2: Linear Algebra, MIT Press (Goodfellow, Bengio and Courville) · read 27 Sept 2026
  3. docsMatrix Methods in Data Analysis, Signal Processing, and Machine Learning (18.065), MIT OpenCourseWare · read 27 Sept 2026
  4. docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
  5. docsLinear (PyTorch 2.14 documentation), PyTorch · read 27 Sept 2026
  6. paperAttention Is All You Need, Vaswani et al., NeurIPS 2017 · read 27 Sept 2026
  7. docsLinear algebra (numpy.linalg), NumPy · read 27 Sept 2026
  8. docstorch.linalg (PyTorch 2.14 documentation), PyTorch · read 27 Sept 2026
  9. docssklearn.decomposition.PCA, scikit-learn · read 27 Sept 2026